Rejected identifiers containing invisible characters, such as
soft hyphens and byte order marks, in is_scholid() and
classify_scholid(). normalize_scholid(),
detect_scholid_type(), and extract_scholid()
now remove those characters before normalizing or matching.
Fixed detect_scholid_type() reporting bare 8-digit
PMIDs such as 29456894 as issn when their
digits happened to pass the ISSN checksum. Bare compact strings are now
detected as ISSN only with a hyphen (2434-561X) or an
ISSN label, so a bare 2434561X is no longer
detected. normalize_scholid(x, "issn") is
unchanged.
classify_scholid(),
detect_scholid_type(), and extract_scholid()
by checking whole vectors instead of one string at a time.The package now supports 20 identifier types (up from 7 in 0.1.1).
Each type provides structural validation, normalization from URLs and
labels, and extraction from free text via the existing
is_scholid(), normalize_scholid(),
extract_scholid(), classify_scholid(), and
detect_scholid_type() APIs.
New types in this release:
W,
A, S, …)orcid)ark:/NAAN/Name)SRR, SRX, SRP, …)GSE,
GSM, GPL, GDS)PRJNA, PRJEB, …)GCA_, GCF_, versioned)Identifier definitions and validation rules are documented in the
scholid_definitions vignette.
is_<type>,
normalize_<type>,
extract_<type>).classify_scholid() and
detect_scholid_type() to avoid redundant work when
resolving types.Initial release.