Triple

T10129597
Position Surface form Disambiguated ID Type / Status
Subject Basukuma E226301 entity
Predicate language P15 FINISHED
Object Sukuma language E181669 NE FINISHED

How this triple was built (2 steps)

Every LLM step that produced this triple, in pipeline order — named-entity classification, the disambiguation choices (the exact options shown, with the pick highlighted), and the generated description. The batch + timestamp of each is in the Provenance table below.

NER Named-entity recognition gpt-5-mini
Instruction
Given a phrase, classify it is english named entity (e.g., persons, organizations, works of art) in Latin script, or not (e.g., literals, dates, URLs, verbose phrases). For disambiguation, the statement where the phrase occurs as object is also given. Please return a JSON object with `phrase` (string, the phrase being analyzed) and `is_ne` (boolean, indicating whether the phrase is a Named Entity).
Input
Phrase: Sukuma language | Statement: [Basukuma, language, Sukuma language]
NED1 Entity disambiguation (via context triple) gpt-5-mini-2025-08-07
Target entity: Sukuma language
Context triple: [Basukuma, language, Sukuma language]
  • A. Sukuma language chosen
    The Sukuma language is a Bantu language spoken primarily by the Sukuma people in northwestern Tanzania.
  • B. Kisukuma language
    Kisukuma is a major Bantu language spoken primarily by the Sukuma people in northwestern Tanzania.
  • C. Nyaturu language
    The Nyaturu language is a Bantu language spoken primarily by the Nyaturu people in central Tanzania.
  • D. Suma language
    The Suma language is a lesser-known Gbaya language spoken by the Suma people in parts of Central Africa.
  • E. Kamba language
    Kamba language is a Bantu language spoken primarily by the Kamba people of Kenya, known for its rich oral traditions and close linguistic ties to other Central Kenya Bantu languages.
  • F. None of above.
  • G. Unsure - the case is ambiguous/there is not enough information to decide.

Provenance (3 batches)

The batch behind each pipeline step, in order, with when it ran. Timestamps are batch-level — stages were processed in waves, so the object chain (NER → NED1 → NEDg → NED2) reads in order, but predicate / elicitation batches can sit in a different wave.

Step Stage Batch ID Status When
creating Elicitation batch_69ca843057b48190a86730167f5d6b98 completed March 30, 2026, 2:09 p.m.
NER Named-entity recognition batch_69cdd333186c819088bbf617967f24fa completed April 2, 2026, 2:23 a.m.
NED1 Entity disambiguation (via context triple) batch_69d2cc7c50b08190a04aa2f58a6c300a completed April 5, 2026, 8:56 p.m.
Created at: March 30, 2026, 9:05 p.m.