Triple

T23359714
Position Surface form Disambiguated ID Type / Status
Subject Kumbewaha language E593150 entity
Predicate belongsToMacroArea P29552 FINISHED
Object Papunesia (linguistic macro-area) NE NERFINISHED

How this triple was built (3 steps)

Every LLM step that produced this triple, in pipeline order — named-entity classification, the disambiguation choices (the exact options shown, with the pick highlighted), and the generated description. The batch + timestamp of each is in the Provenance table below.

NER Named-entity recognition gpt-5-mini
Instruction
Given a phrase, classify it is english named entity (e.g., persons, organizations, works of art) in Latin script, or not (e.g., literals, dates, URLs, verbose phrases). For disambiguation, the statement where the phrase occurs as object is also given. Please return a JSON object with `phrase` (string, the phrase being analyzed) and `is_ne` (boolean, indicating whether the phrase is a Named Entity).
Input
Phrase: Papunesia (linguistic macro-area) | Statement: [Kumbewaha language, belongsToMacroArea, Papunesia (linguistic macro-area)]
NED1 Entity disambiguation (via context triple) gpt-5-mini-2025-08-07
Target entity: Papunesia (linguistic macro-area)
Context triple: [Kumbewaha language, belongsToMacroArea, Papunesia (linguistic macro-area)]
  • A. Plateau linguistic area
    The Plateau linguistic area is a region of the North American Plateau where diverse Indigenous languages, often from different families, share common structural features due to long-term contact and interaction.
  • B. Punu languages
    Punu languages are a subgroup of Bantu languages spoken primarily by the Punu people of Gabon and neighboring regions, known for their close linguistic and cultural ties to other Western Bantu groups.
  • C. Torres–Banks languages area
    The Torres–Banks languages area is a region in northern Vanuatu known for its dense cluster of closely related Oceanic languages and rich linguistic diversity.
  • D. Pamir Sprachbund
    The Pamir Sprachbund is a linguistic convergence area in the Pamir Mountains where several genetically distinct languages, including Eastern Iranian and others, have developed shared structural features through long-term contact.
  • E. Guadalcanal linguistic area
    The Guadalcanal linguistic area is a region on the island of Guadalcanal in the Solomon Islands characterized by a cluster of related and interacting Oceanic languages that share common structural and lexical features.
  • F. None of above. chosen
  • G. Unsure - the case is ambiguous/there is not enough information to decide.
NED2 Entity disambiguation (via description) gpt-5-mini-2025-08-07
Target entity: Papunesia (linguistic macro-area)
Target entity description: Papunesia is a linguistic macro-area encompassing the diverse and often unrelated languages spoken across New Guinea, nearby islands, and parts of eastern Indonesia and Melanesia.
  • A. Plateau linguistic area
    The Plateau linguistic area is a region of the North American Plateau where diverse Indigenous languages, often from different families, share common structural features due to long-term contact and interaction.
  • B. Punu languages
    Punu languages are a subgroup of Bantu languages spoken primarily by the Punu people of Gabon and neighboring regions, known for their close linguistic and cultural ties to other Western Bantu groups.
  • C. Torres–Banks languages area
    The Torres–Banks languages area is a region in northern Vanuatu known for its dense cluster of closely related Oceanic languages and rich linguistic diversity.
  • D. Pamir Sprachbund
    The Pamir Sprachbund is a linguistic convergence area in the Pamir Mountains where several genetically distinct languages, including Eastern Iranian and others, have developed shared structural features through long-term contact.
  • E. Guadalcanal linguistic area
    The Guadalcanal linguistic area is a region on the island of Guadalcanal in the Solomon Islands characterized by a cluster of related and interacting Oceanic languages that share common structural and lexical features.
  • F. None of above. chosen

Provenance (2 batches)

The batch behind each pipeline step, in order, with when it ran. Timestamps are batch-level — stages were processed in waves, so the object chain (NER → NED1 → NEDg → NED2) reads in order, but predicate / elicitation batches can sit in a different wave.

Step Stage Batch ID Status When
creating Elicitation batch_69e25d24d2a4819092e6ede74c2a918d completed April 17, 2026, 4:17 p.m.
NER Named-entity recognition batch_69f19a1a39988190b4b4993b80d5a5f6 completed April 29, 2026, 5:41 a.m.
Created at: April 17, 2026, 5:30 p.m.