Triple
T23473307
| Position | Surface form | Disambiguated ID | Type / Status |
|---|---|---|---|
| Subject | Tomia dialect |
E570185
|
entity |
| Predicate | belongsToDialectGroup |
P1254
|
FINISHED |
| Object | Tukang Besi dialect continuum |
—
|
NE NERFINISHED |
How this triple was built (3 steps)
Every LLM step that produced this triple, in pipeline order — named-entity classification, the disambiguation choices (the exact options shown, with the pick highlighted), and the generated description. The batch + timestamp of each is in the Provenance table below.
NER
Named-entity recognition
gpt-5-mini
Instruction
Given a phrase, classify it is english named entity (e.g., persons, organizations, works of art) in Latin script, or not (e.g., literals, dates, URLs, verbose phrases). For disambiguation, the statement where the phrase occurs as object is also given. Please return a JSON object with `phrase` (string, the phrase being analyzed) and `is_ne` (boolean, indicating whether the phrase is a Named Entity).
Input
Phrase: Tukang Besi dialect continuum | Statement: [Tomia dialect, belongsToDialectGroup, Tukang Besi dialect continuum]
NED1
Entity disambiguation (via context triple)
gpt-5-mini-2025-08-07
Target entity: Tukang Besi dialect continuum Context triple: [Tomia dialect, belongsToDialectGroup, Tukang Besi dialect continuum]
-
A.
Tat dialect continuum
The Tat dialect continuum is a group of closely related Southwestern Iranian dialects spoken primarily in the eastern Caucasus region, notably in parts of Azerbaijan and Russia.
-
B.
Sasak dialect continuum
The Sasak dialect continuum is a group of closely related Austronesian speech varieties spoken by the Sasak people on the island of Lombok in Indonesia, exhibiting gradual linguistic variation across regions rather than sharply defined dialect boundaries.
-
C.
Fala dialect continuum
The Fala dialect continuum is a group of closely related Romance varieties spoken in a small area of Extremadura in western Spain, often considered transitional between Galician-Portuguese and Spanish.
-
D.
Tabasaran dialect continuum
The Tabasaran dialect continuum is a group of closely related Northeast Caucasian speech varieties of the Tabasaran language, spoken primarily in southern Dagestan, Russia.
-
E.
Javanese dialect continuum
The Javanese dialect continuum is a range of closely related Javanese language varieties spoken across Java whose neighboring dialects are mutually intelligible but show gradual linguistic differences over distance.
- F. None of above. chosen
- G. Unsure - the case is ambiguous/there is not enough information to decide.
NED2
Entity disambiguation (via description)
gpt-5-mini-2025-08-07
Target entity: Tukang Besi dialect continuum Target entity description: The Tukang Besi dialect continuum is a group of closely related Austronesian dialects spoken in the Tukang Besi Islands of Southeast Sulawesi, Indonesia, forming a chain of mutually intelligible local speech varieties.
-
A.
Tat dialect continuum
The Tat dialect continuum is a group of closely related Southwestern Iranian dialects spoken primarily in the eastern Caucasus region, notably in parts of Azerbaijan and Russia.
-
B.
Sasak dialect continuum
The Sasak dialect continuum is a group of closely related Austronesian speech varieties spoken by the Sasak people on the island of Lombok in Indonesia, exhibiting gradual linguistic variation across regions rather than sharply defined dialect boundaries.
-
C.
Fala dialect continuum
The Fala dialect continuum is a group of closely related Romance varieties spoken in a small area of Extremadura in western Spain, often considered transitional between Galician-Portuguese and Spanish.
-
D.
Tabasaran dialect continuum
The Tabasaran dialect continuum is a group of closely related Northeast Caucasian speech varieties of the Tabasaran language, spoken primarily in southern Dagestan, Russia.
-
E.
Javanese dialect continuum
The Javanese dialect continuum is a range of closely related Javanese language varieties spoken across Java whose neighboring dialects are mutually intelligible but show gradual linguistic differences over distance.
- F. None of above. chosen
Provenance (2 batches)
The batch behind each pipeline step, in order, with when it ran. Timestamps are batch-level — stages were processed in waves, so the object chain (NER → NED1 → NEDg → NED2) reads in order, but predicate / elicitation batches can sit in a different wave.
| Step | Stage | Batch ID | Status | When |
|---|---|---|---|---|
| creating | Elicitation | batch_69e245af8a88819084f2704f6d265a92 |
completed | April 17, 2026, 2:37 p.m. |
| NER | Named-entity recognition | batch_69f1a70244208190bbd8f58ac16d4399 |
completed | April 29, 2026, 6:36 a.m. |
Created at: April 17, 2026, 5:58 p.m.