Triple

T14815582
Position Surface form Disambiguated ID Type / Status
Subject Embu people E348305 entity
Predicate language P15 FINISHED
Object Embu language E267656 NE FINISHED

How this triple was built (2 steps)

Every LLM step that produced this triple, in pipeline order — named-entity classification, the disambiguation choices (the exact options shown, with the pick highlighted), and the generated description. The batch + timestamp of each is in the Provenance table below.

NER Named-entity recognition gpt-5-mini
Instruction
Given a phrase, classify it is english named entity (e.g., persons, organizations, works of art) in Latin script, or not (e.g., literals, dates, URLs, verbose phrases). For disambiguation, the statement where the phrase occurs as object is also given. Please return a JSON object with `phrase` (string, the phrase being analyzed) and `is_ne` (boolean, indicating whether the phrase is a Named Entity).
Input
Phrase: Embu language | Statement: [Embu people, language, Embu language]
NED1 Entity disambiguation (via context triple) gpt-5-mini-2025-08-07
Target entity: Embu language
Context triple: [Embu people, language, Embu language]
  • A. Embu language chosen
    The Embu language is a Bantu language spoken primarily by the Embu people of central Kenya, closely related to other Mount Kenya languages such as Kikuyu and Meru.
  • B. Embera Katío language
    Embera Katío language is an indigenous Chocoan language of Colombia spoken by the Embera Katío people, known for its rich oral tradition and endangered status.
  • C. Apurinã language
    The Apurinã language is an indigenous Arawakan language spoken by the Apurinã people of the Brazilian Amazon, known for its complex verbal morphology and endangered status.
  • D. Tapirapé language
    Tapirapé is an indigenous Tupian language spoken by the Tapirapé people of Brazil, known for its complex morphology and endangered status.
  • E. Behoa language
    The Behoa language is an Austronesian language spoken in Central Sulawesi, Indonesia, known for its place within the Badaic subgroup of the Kaili–Pamona languages.
  • F. None of above.
  • G. Unsure - the case is ambiguous/there is not enough information to decide.

Provenance (3 batches)

The batch behind each pipeline step, in order, with when it ran. Timestamps are batch-level — stages were processed in waves, so the object chain (NER → NED1 → NEDg → NED2) reads in order, but predicate / elicitation batches can sit in a different wave.

Step Stage Batch ID Status When
creating Elicitation batch_69d822eb8f588190bf53445e730a934f completed April 9, 2026, 10:06 p.m.
NER Named-entity recognition batch_69decfe0e89c81908c0e1fe2bc3ebcfc completed April 14, 2026, 11:38 p.m.
NED1 Entity disambiguation (via context triple) batch_69fe389598848190ba15e6eea2ba2903 completed May 8, 2026, 7:25 p.m.
Created at: April 10, 2026, 1:49 a.m.