Triple

T23473307
Position Surface form Disambiguated ID Type / Status
Subject Tomia dialect E570185 entity
Predicate belongsToDialectGroup P1254 FINISHED
Object Tukang Besi dialect continuum NE NERFINISHED

How this triple was built (3 steps)

Every LLM step that produced this triple, in pipeline order — named-entity classification, the disambiguation choices (the exact options shown, with the pick highlighted), and the generated description. The batch + timestamp of each is in the Provenance table below.

NER Named-entity recognition gpt-5-mini
Instruction
Given a phrase, classify it is english named entity (e.g., persons, organizations, works of art) in Latin script, or not (e.g., literals, dates, URLs, verbose phrases). For disambiguation, the statement where the phrase occurs as object is also given. Please return a JSON object with `phrase` (string, the phrase being analyzed) and `is_ne` (boolean, indicating whether the phrase is a Named Entity).
Input
Phrase: Tukang Besi dialect continuum | Statement: [Tomia dialect, belongsToDialectGroup, Tukang Besi dialect continuum]
NED1 Entity disambiguation (via context triple) gpt-5-mini-2025-08-07
Target entity: Tukang Besi dialect continuum
Context triple: [Tomia dialect, belongsToDialectGroup, Tukang Besi dialect continuum]
  • A. Tat dialect continuum
    The Tat dialect continuum is a group of closely related Southwestern Iranian dialects spoken primarily in the eastern Caucasus region, notably in parts of Azerbaijan and Russia.
  • B. Sasak dialect continuum
    The Sasak dialect continuum is a group of closely related Austronesian speech varieties spoken by the Sasak people on the island of Lombok in Indonesia, exhibiting gradual linguistic variation across regions rather than sharply defined dialect boundaries.
  • C. Fala dialect continuum
    The Fala dialect continuum is a group of closely related Romance varieties spoken in a small area of Extremadura in western Spain, often considered transitional between Galician-Portuguese and Spanish.
  • D. Tabasaran dialect continuum
    The Tabasaran dialect continuum is a group of closely related Northeast Caucasian speech varieties of the Tabasaran language, spoken primarily in southern Dagestan, Russia.
  • E. Javanese dialect continuum
    The Javanese dialect continuum is a range of closely related Javanese language varieties spoken across Java whose neighboring dialects are mutually intelligible but show gradual linguistic differences over distance.
  • F. None of above. chosen
  • G. Unsure - the case is ambiguous/there is not enough information to decide.
NED2 Entity disambiguation (via description) gpt-5-mini-2025-08-07
Target entity: Tukang Besi dialect continuum
Target entity description: The Tukang Besi dialect continuum is a group of closely related Austronesian dialects spoken in the Tukang Besi Islands of Southeast Sulawesi, Indonesia, forming a chain of mutually intelligible local speech varieties.
  • A. Tat dialect continuum
    The Tat dialect continuum is a group of closely related Southwestern Iranian dialects spoken primarily in the eastern Caucasus region, notably in parts of Azerbaijan and Russia.
  • B. Sasak dialect continuum
    The Sasak dialect continuum is a group of closely related Austronesian speech varieties spoken by the Sasak people on the island of Lombok in Indonesia, exhibiting gradual linguistic variation across regions rather than sharply defined dialect boundaries.
  • C. Fala dialect continuum
    The Fala dialect continuum is a group of closely related Romance varieties spoken in a small area of Extremadura in western Spain, often considered transitional between Galician-Portuguese and Spanish.
  • D. Tabasaran dialect continuum
    The Tabasaran dialect continuum is a group of closely related Northeast Caucasian speech varieties of the Tabasaran language, spoken primarily in southern Dagestan, Russia.
  • E. Javanese dialect continuum
    The Javanese dialect continuum is a range of closely related Javanese language varieties spoken across Java whose neighboring dialects are mutually intelligible but show gradual linguistic differences over distance.
  • F. None of above. chosen

Provenance (2 batches)

The batch behind each pipeline step, in order, with when it ran. Timestamps are batch-level — stages were processed in waves, so the object chain (NER → NED1 → NEDg → NED2) reads in order, but predicate / elicitation batches can sit in a different wave.

Step Stage Batch ID Status When
creating Elicitation batch_69e245af8a88819084f2704f6d265a92 completed April 17, 2026, 2:37 p.m.
NER Named-entity recognition batch_69f1a70244208190bbd8f58ac16d4399 completed April 29, 2026, 6:36 a.m.
Created at: April 17, 2026, 5:58 p.m.