Triple

T21603527
Position Surface form Disambiguated ID Type / Status
Subject Papuan languages of the Solomon Islands E533110 entity
Predicate languageFamilyProposal P138721 FINISHED
Object East Papuan hypothesis NE NERFINISHED

How this triple was built (3 steps)

Every LLM step that produced this triple, in pipeline order — named-entity classification, the disambiguation choices (the exact options shown, with the pick highlighted), and the generated description. The batch + timestamp of each is in the Provenance table below.

NER Named-entity recognition gpt-5-mini
Instruction
Given a phrase, classify it is english named entity (e.g., persons, organizations, works of art) in Latin script, or not (e.g., literals, dates, URLs, verbose phrases). For disambiguation, the statement where the phrase occurs as object is also given. Please return a JSON object with `phrase` (string, the phrase being analyzed) and `is_ne` (boolean, indicating whether the phrase is a Named Entity).
Input
Phrase: East Papuan hypothesis | Statement: [Papuan languages of the Solomon Islands, languageFamilyProposal, East Papuan hypothesis]
NED1 Entity disambiguation (via context triple) gpt-5-mini-2025-08-07
Target entity: East Papuan hypothesis
Context triple: [Papuan languages of the Solomon Islands, languageFamilyProposal, East Papuan hypothesis]
  • A. Trans-New Guinea hypothesis
    The Trans-New Guinea hypothesis is a major linguistic proposal that posits a large family of related Papuan languages spread across much of New Guinea and nearby islands.
  • B. Sunda-Sulawesi hypothesis
    The Sunda-Sulawesi hypothesis is a proposed subgrouping within the Austronesian language family that suggests a closer genetic relationship among certain languages spoken in western Indonesia and surrounding regions.
  • C. Greater North Borneo hypothesis
    The Greater North Borneo hypothesis is a linguistic proposal that groups several Austronesian languages of Borneo and surrounding regions into a single higher-order subgroup based on shared innovations and historical relationships.
  • D. Austric hypothesis
    The Austric hypothesis is a proposed but controversial macro-family theory suggesting a common origin for several language families of Southeast Asia and the Pacific, including Austroasiatic and Austronesian.
  • E. Yok-Utian hypothesis
    The Yok-Utian hypothesis is a proposed linguistic theory suggesting a genetic relationship between the Yokuts and Utian language families of California.
  • F. None of above. chosen
  • G. Unsure - the case is ambiguous/there is not enough information to decide.
NED2 Entity disambiguation (via description) gpt-5-mini-2025-08-07
Target entity: East Papuan hypothesis
Target entity description: The East Papuan hypothesis is a proposed grouping of several non-Austronesian languages of island Melanesia that suggests they share a common ancestral origin distinct from other Papuan language families.
  • A. Trans-New Guinea hypothesis
    The Trans-New Guinea hypothesis is a major linguistic proposal that posits a large family of related Papuan languages spread across much of New Guinea and nearby islands.
  • B. Sunda-Sulawesi hypothesis
    The Sunda-Sulawesi hypothesis is a proposed subgrouping within the Austronesian language family that suggests a closer genetic relationship among certain languages spoken in western Indonesia and surrounding regions.
  • C. Greater North Borneo hypothesis
    The Greater North Borneo hypothesis is a linguistic proposal that groups several Austronesian languages of Borneo and surrounding regions into a single higher-order subgroup based on shared innovations and historical relationships.
  • D. Austric hypothesis
    The Austric hypothesis is a proposed but controversial macro-family theory suggesting a common origin for several language families of Southeast Asia and the Pacific, including Austroasiatic and Austronesian.
  • E. Yok-Utian hypothesis
    The Yok-Utian hypothesis is a proposed linguistic theory suggesting a genetic relationship between the Yokuts and Utian language families of California.
  • F. None of above. chosen

Provenance (2 batches)

The batch behind each pipeline step, in order, with when it ran. Timestamps are batch-level — stages were processed in waves, so the object chain (NER → NED1 → NEDg → NED2) reads in order, but predicate / elicitation batches can sit in a different wave.

Step Stage Batch ID Status When
creating Elicitation batch_69e0c46364608190a337dc8720dc2a35 completed April 16, 2026, 11:13 a.m.
NER Named-entity recognition batch_69ef17e31c80819090fdd63b5c103acb completed April 27, 2026, 8:01 a.m.
Created at: April 16, 2026, 6:33 p.m.