Draft. This essay is an unreviewed draft. Its sources have not been checked by a named person and no domain reviewer has approved it. Treat every claim as provisional.
Central question. How do the terminology foundation, source registry, and concept composer fit together?
Key points
- NucLex is planned as three connected functions: a nuclear medicine niche anchored in SNOMED CT, a curated registry of preferred external authorities, and a composer that builds compound concepts from existing parts.
- The first rule is to extend rather than duplicate: reference what SNOMED CT or a recognized authority already defines, and create only what is missing.
- Eight sources were named in the founding discussion, each for a different kind of concept: SNOMED CT, PubChem, FDA GSRS, RxNorm, UniProt, NCI Thesaurus, LOINC, and RadLex.
- A composed concept, such as a radiolabeled ligand conjugate, would be assembled from a radionuclide, a ligand, and a target, with a provenance record naming each part and its source.
- Nothing composed by software would become official without human review. Connectors, endpoints, mapping stores, and promotion are future capabilities, described in plan language.
A niche, not a rival
Nuclear medicine and theranostics occupy an awkward position in medical terminology. The field is small enough that general clinical vocabularies cover it unevenly, and technical enough that its concepts are shared across chemistry, physics, pharmacology, and oncology rather than owned by any one of them. A single Radiopharmaceutical is at once a chemical substance, a regulated drug product, a radiation source, a molecular probe for a receptor, and the subject of a clinical procedure. Each facet has a community with a vocabulary, and none was designed with the others in mind.
The temptation is to build a complete vocabulary from the ground up. The founding discussion of NucLex rejected that path at the outset (R§1): the requirement was a nuclear medicine ontology and terminology service anchored in SNOMED CT that would complement it rather than re-create it. SNOMED CT is the most widely adopted clinical terminology in electronic health records, it has a defined concept model and editorial rules, and it already carries much of what a department records: diagnoses, procedures, body structures, findings, and substances. A niche vocabulary that ignored it would cut nuclear medicine off from the record systems it has to live in.
The niche is therefore defined by subtraction. NucLex concerns itself with what SNOMED CT and the other authorities do not express well enough for nuclear oncology: the fine-grained identity of radiopharmaceuticals (nuclide, chemical form, chelator, ligand, formulation), the theranostic pairing of a diagnostic and a therapeutic agent, the relationship between a Ligand and its Molecular target, and, in later stages, the interactions between molecules that a computational model needs (see Interactional ontology). The monograph SNOMED CT and the NucLex niche develops that boundary. This essay concerns the three functions meant to operate inside it.
Extension versus duplication
The founding discussion also set a rule of conduct: "avoid reinventing the wheel" (R§3). Existing concepts from other ontologies and terminologies should be brought into NucLex rather than re-created there. The example offered was the United States Food and Drug Administration substance registry: a compound with a registered identity should be referenced by that identity, not given a fresh definition that would drift from the original.
The distinction is easy to honor in principle and violate in practice. Duplication happens when a project creates a new identifier and definition for something an authoritative source already defines. Copies diverge: the source corrects an error or retires the concept, and the copy does not follow. Two systems that believed they meant the same thing now disagree, and nobody can tell which is right without a mapping that nobody maintained. Every duplicate also needs a curator, and curation spread across many small vocabularies is mostly not done.
Extension happens when a project adds a concept the source lacks, places it in a defined relationship to concepts the source has, and leaves the source's own concepts alone. In SNOMED CT this is the formal mechanism of an extension: a namespace in which an organization authors concepts additional to the International Edition, modeled under the same editorial rules and distinguishable by identifier from the core. The founding discussion proposed that the NucLex niche take this form (R§1 to 2).
The practical test is a question asked before any concept is authored: does something with this meaning already exist in a source we recognize? If so, NucLex references it. If a close but not identical concept exists, NucLex records the relationship (broader, narrower, or related) rather than pretending equivalence. Only when the answer is no does authoring begin. The governance discussion later named this the reuse check (R§6). In W0 it is editorial discipline; in the planned platform it would be a step the tooling enforces.
A Concept page in W0 may list an external authority's home page as further reading, but it asserts no formal mapping to any identifier there. Formal identifiers are never guessed, and the editorial IDs on this site are local labels that claim no membership in any released ontology.
The curated registry: which authority for which concept
The first pillar fixes the clinical foundation. The second answers a question that arises as soon as a radiopharmaceutical concept is taken apart: when a component is not clinical at all, where should its identity come from? A protein is not a SNOMED CT concern in the way a diagnosis is; a chemical structure is not a medication. The founding discussion described this pillar as a curated shelf (R§4): NucLex should know which external source to prefer for which domain, point to it, and say so. The sources named, and the role each was assigned, follow. The roles are intended usage, not a claim about the full scope of any source.
| Source | Issuing body | Intended role in NucLex |
|---|---|---|
| SNOMED CT | SNOMED International | Clinical foundation: diagnoses, procedures, findings, body structures, substances; home of the planned extension |
| PubChem | NCBI, U.S. National Library of Medicine | Chemical structures and identity: small molecules, chelators, precursors |
| FDA GSRS | U.S. FDA with NCATS | Regulated substance identity, including the Unique Ingredient Identifier (UNII) |
| RxNorm | U.S. National Library of Medicine | Medication terminology: clinical drugs as prescribed and dispensed in the United States |
| UniProt | UniProt Consortium | Protein concepts: sequence, function, annotation, including receptor targets |
| NCI Thesaurus | National Cancer Institute | Cancer concepts: neoplasms, biomarkers, agents, trial terminology |
| LOINC | Regenstrief Institute | Observation and measurement codes: laboratory tests and clinical observations |
| RadLex | Radiological Society of North America | Radiology terminology: imaging anatomy, findings, procedure descriptors |
Three features matter more than the list. First, each source is preferred for a kind of concept, not for everything. PubChem and GSRS both describe chemicals, but PubChem answers "what is this structure" while GSRS answers "what is this substance as the regulator recognizes it." A radiochemist and a regulatory specialist need both answers, and the registry should say which is which rather than merge them. Second, the registry is itself curated with provenance: an entry would record why a source is preferred, for which concept kinds, which version was consulted, and who decided. Third, the registry is planned, not built. The discussion proposed a source authority registry, term resolvers, a mapping store in the style of the Simple Standard for Sharing Ontological Mappings (SSSOM), and a composer interface as Version 2 additions (R§4 to 5), and is explicit that no connectors had been built. None exists in W0.
The composer: a hypothetical compound concept
The third pillar distinguishes NucLex from a well-organized bookmark list. The founding discussion called it an "ontology studio" or "synthesizer" (R§4): a function that combines concepts from different sources into more complex concepts, each part traceable to its origin. Nuclear medicine needs this because radiopharmaceuticals and theranostic pairs are compound by nature: their meaning is the meaning of their parts plus the way the parts are joined.
What follows is a synthetic example. The names and identifiers are invented for teaching and correspond to no real agent, no ontology record, and no entry in any source above. Teaching identifiers carry the prefix TEACH to keep them apart from NucLex editorial IDs (NX-C) and from any external identifier.
Suppose a radiochemist wants a single concept for a therapeutic agent: a beta-emitting Radionuclide attached, through a chelator, to a small peptide that binds a receptor over-expressed on tumor cells. In most records today that agent is free text, a drug code silent about the nuclide, or a substance code silent about the target. The composer would build it from three parts.
The radionuclide. TEACH-RN-01, "Nuclide X." Its definition (element, mass number, decay mode, half-life) belongs to physics. In the planned registry its identity would be anchored to a substance record and, where SNOMED CT has a matching substance concept, to that concept.
The ligand. TEACH-LG-01, "Peptide Y." A short synthetic peptide with a chelator moiety. Its structure would be anchored to PubChem and its regulated identity to GSRS, if and when such records exist.
The target. TEACH-TG-01, "Receptor Z." A cell-surface receptor. Its identity as a protein would be anchored to UniProt and its role as a cancer biomarker to NCI Thesaurus. No external identifier is asserted for any of the three.
The composed concept, TEACH-RP-01, "Nuclide X-labeled Peptide Y targeting Receptor Z," is not a fourth independent definition. It is a structured expression: this is a radiopharmaceutical; its radionuclide is TEACH-RN-01; its ligand is TEACH-LG-01; its intended molecular target is TEACH-TG-01; the linkage is chelation. In the planned platform the expression would be written in a form a terminology server can evaluate, as SNOMED CT post-coordinated expressions combine existing concepts with defined attributes. The syntax is an implementation decision for Technical V1 and later.
The Provenance record that would travel with TEACH-RP-01 is the point of the exercise. Shown as a table, and labeled a synthetic example:
| Field | Value in the example |
|---|---|
| Composed concept | TEACH-RP-01 |
| Component: radionuclide | TEACH-RN-01, authority per registry entry (not resolved in W0) |
| Component: ligand | TEACH-LG-01, authorities per registry entry (not resolved in W0) |
| Component: target | TEACH-TG-01, authorities per registry entry (not resolved in W0) |
| Composition pattern | radiopharmaceutical: radionuclide, ligand, molecular target, chelation linkage |
| Proposed by | a named person, or an agent with model and prompt version logged |
| Evidence | supporting literature or documentation (none for a teaching example) |
| Source versions consulted | recorded per authority at composition time |
| Review state | candidate |
Every line answers a question a later reader will ask: which nuclide, which peptide, which receptor, according to whom, on what evidence, and has anyone checked? A concept that cannot answer those questions is a label. A concept that can is knowledge someone else can audit, reuse, and correct.
Review before promotion
The last line of that table, review state, is where the composer meets governance. The founding discussion held one boundary constant from the first platform proposal through the later agentic addition: software and agents may propose, and humans approve (R§4, R§10). A composed expression would become official only after formal review (R§5). The monograph Human review of machine proposals examines why that boundary is not a formality; Human review gives the short form.
In plan language, the lifecycle of TEACH-RP-01 would run as follows. A candidate is created, by a person at a composer interface or by an agent extracting terms from literature, and enters a queue with its provenance record and evidence. A reviewer with domain competence (a radiochemist or nuclear medicine physician) checks that the parts are the right parts, that no existing concept already expresses the whole (the reuse check), and that the composition is modeled correctly against the editorial rules of the niche. The reviewer accepts, returns with comments, or rejects. An accepted candidate is promoted: it receives a stable identifier in the extension namespace, its mappings to external authorities are recorded in the mapping store with evidence and confidence, and it is published in a versioned release. A rejected candidate is retained with its history.
Later stages add weight without changing the shape. The Version 3 discussion proposed proposal, triage, authoring, internal and expert review, public comment, reconciliation, approval, publication, and maintenance including deprecation (R§5 to 6). The gate stays constant: nothing composed by machine is official until a named person has said so, and that decision is part of the record.
Two clarifications keep this honest. In W0 none of this runs in software; the present workflow is editorial (brief, draft, source check, domain review, editorial approval). And review does not make a composed concept true in the world; it makes the concept well-formed and its provenance complete. Whether Peptide Y really binds Receptor Z with useful selectivity is a question for pharmacology and clinical evidence, and the planned interactional layer is where such claims would be represented, with their uncertainty, rather than assumed.
Why three and not one
The three functions cannot be collapsed, because each answers a different failure. Without the first, NucLex would float free of the clinical record. Without the second, it would have to define chemistry, proteins, and measurements itself, and would do so worse than the bodies whose job that is. Without the third, it could only point at parts; it could not say what a radiopharmaceutical is in a way a terminology server or a model could use. The discussion's later note that editorial content and the ontology graph should stay distinct while linked (R§34) applies: essays and monographs explain, concept pages orient, and the planned graph is where composed knowledge would live with its identifiers, mappings, and versions.
The pillars are also an order of work. Reuse is cheaper than curation, and curation is cheaper than composition. NucLex begins with the publication you are reading because that discipline has to be learned in prose before it can be enforced in software.
Limitations
This essay describes intended behavior recovered from a design discussion and from official documentation of the named sources, not a working system. No connector, terminology endpoint, mapping store, composer interface, or promotion workflow exists in W0, and the roles in the registry table are editorial intentions that may change once licensing and update practices are examined.
The compound concept is a teaching construction with local identifiers and fictitious parts, not validated against SNOMED CT editorial rules or any source's data model. How SNOMED International itself authors and approves new concepts was asked in the founding discussion and not answered (R§5); that research remains open. The essay has not been source-checked; the sources below are cited for the existence and general role of each resource, not for any identifier or data element.
Sources and further reading
Official documentation of the named sources. URLs are references to the issuing bodies, not mappings.
- SNOMED International. SNOMED CT. https://www.snomed.org. Editorial Guide: https://confluence.ihtsdotools.org/display/DOCEG
- National Center for Biotechnology Information, U.S. National Library of Medicine. PubChem. https://pubchem.ncbi.nlm.nih.gov
- U.S. Food and Drug Administration and National Center for Advancing Translational Sciences. Global Substance Registration System (GSRS). https://gsrs.ncats.nih.gov
- U.S. National Library of Medicine. RxNorm. https://www.nlm.nih.gov/research/umls/rxnorm/index.html
- UniProt Consortium. UniProt. https://www.uniprot.org
- National Cancer Institute, Enterprise Vocabulary Services. NCI Thesaurus. https://ncithesaurus.nci.nih.gov
- Regenstrief Institute. LOINC. https://loinc.org
- Radiological Society of North America. RadLex. https://radlex.org
- Matentzoglu N, et al. 2022. A Simple Standard for Sharing Ontological Mappings (SSSOM). Database (Oxford). Specification: https://mapping-commons.github.io/sssom/
- Smith B, et al. 2007. The OBO Foundry: coordinated evolution of ontologies to support biomedical data integration. Nature Biotechnology. Principles: https://obofoundry.org/principles/fp-000-summary.html
- NucLex Detailed Discussion Record, sections 1 to 10 and 34 (project document; basis for statements about intended behavior).
On this site: From living publication to interactional knowledge, Concept map, Identifier.
Source list as recorded in the manuscript metadata (12)
- SNOMED International. SNOMED CT Editorial Guide. https://confluence.ihtsdotools.org/display/DOCEG
- SNOMED International. SNOMED CT home page. https://www.snomed.org
- National Center for Biotechnology Information (NCBI), U.S. National Library of Medicine. PubChem. https://pubchem.ncbi.nlm.nih.gov
- U.S. Food and Drug Administration and National Center for Advancing Translational Sciences. Global Substance Registration System (GSRS). https://gsrs.ncats.nih.gov
- U.S. National Library of Medicine. RxNorm. https://www.nlm.nih.gov/research/umls/rxnorm/index.html
- UniProt Consortium. UniProt. https://www.uniprot.org
- National Cancer Institute, Enterprise Vocabulary Services. NCI Thesaurus. https://ncithesaurus.nci.nih.gov
- Regenstrief Institute. LOINC. https://loinc.org
- Radiological Society of North America. RadLex. https://radlex.org
- Matentzoglu N, et al. 2022. A Simple Standard for Sharing Ontological Mappings (SSSOM). Database (Oxford). Specification: https://mapping-commons.github.io/sssom/
- Smith B, et al. 2007. The OBO Foundry: coordinated evolution of ontologies to support biomedical data integration. Nature Biotechnology. Principles: https://obofoundry.org/principles/fp-000-summary.html
- HL7 International. FHIR Terminology Module. https://hl7.org/fhir/terminology-module.html
These citations have not yet been verified by a named source checker. A citation existing is not the same as a citation supporting the precise claim.