Research & methodology
Measure meaning before claiming universality.
The project treats multilingual semantic alignment as an empirical retrieval problem with explicit failure cases, not as proof that language-specific distinctions have disappeared.
01 · Objective
Concept resolution, not generic similarity
The target task is to identify which registry concept best represents an expression. Topic similarity is useful evidence, but it is not equivalent to semantic identity.
02 · Training
Cross-language positives plus hard negatives
Equivalent expressions across languages should converge. Closely related but incorrect concepts should remain distinguishable. A useful training corpus therefore contains reviewed translations, paraphrases, definitions, and hard negatives.
03 · Evaluation
A benchmark that rewards discrimination
- Concept Recall@1 and Recall@5
- Mean reciprocal rank
- Cross-language concept agreement
- False-neighbor rate
- Centroid and language-family spread
- Top-1 / top-2 score margin
- Abstention precision and recall
04 · Abstention
Unknown is a valid result
A resolver should decline to assign a concept when its calibrated evidence is weak or competing candidates are too close. Forced classification is treated as an error mode.
05 · Provenance
Every semantic claim should be traceable
Concept definitions, expressions, renderings, relationships, model profiles, and evaluation outcomes are versioned independently so changes can be inspected over time.