Triple

T7192483
Position Surface form Disambiguated ID Type / Status
Subject Newcomb–Benford law E167726 entity
Predicate relatedConcept P37 FINISHED
Object Zipf's law E595317 NE FINISHED

How this triple was built (2 steps)

Every LLM step that produced this triple, in pipeline order — named-entity classification, the disambiguation choices (the exact options shown, with the pick highlighted), and the generated description. The batch + timestamp of each is in the Provenance table below.

NER Named-entity recognition gpt-5-mini
Instruction
Given a phrase, classify it is english named entity (e.g., persons, organizations, works of art) in Latin script, or not (e.g., literals, dates, URLs, verbose phrases). For disambiguation, the statement where the phrase occurs as object is also given. Please return a JSON object with `phrase` (string, the phrase being analyzed) and `is_ne` (boolean, indicating whether the phrase is a Named Entity).
Input
Phrase: Zipf's law | Statement: [Newcomb–Benford law, relatedConcept, Zipf's law]
NED1 Entity disambiguation (via context triple) gpt-5-mini-2025-08-07
Target entity: Zipf's law
Context triple: [Newcomb–Benford law, relatedConcept, Zipf's law]
  • A. Zipf's law chosen
    Zipf's law is an empirical statistical principle observing that in many datasets, such as word frequencies in natural language, the frequency of an item is inversely proportional to its rank in a frequency table.
  • B. Szemerényi's law
    Szemerényi's law is a sound law in Proto-Indo-European linguistics that explains the loss of certain final consonants with compensatory lengthening of the preceding vowel.
  • C. Zipf
    Zipf is a surname most notably associated with linguist George Kingsley Zipf, known for formulating Zipf's law about word frequency distributions.
  • D. Newcomb–Benford law
    The Newcomb–Benford law is a statistical principle stating that in many naturally occurring datasets, the leading digits are distributed logarithmically, with smaller digits (especially 1) appearing as the first digit more frequently than larger ones.
  • E. Pareto distribution
    The Pareto distribution is a power-law probability distribution often used to model phenomena with heavy tails and strong inequality, such as wealth or city sizes.
  • F. None of above.
  • G. Unsure - the case is ambiguous/there is not enough information to decide.

Provenance (3 batches)

The batch behind each pipeline step, in order, with when it ran. Timestamps are batch-level — stages were processed in waves, so the object chain (NER → NED1 → NEDg → NED2) reads in order, but predicate / elicitation batches can sit in a different wave.

Step Stage Batch ID Status When
creating Elicitation batch_69c6888b5248819090499a884ee3ec39 completed March 27, 2026, 1:39 p.m.
NER Named-entity recognition batch_69c6e901ea1481908a9e44f96dd4b553 completed March 27, 2026, 8:30 p.m.
NED1 Entity disambiguation (via context triple) batch_69c7bf95a1a0819099d252f037c318c7 completed March 28, 2026, 11:46 a.m.
Created at: March 27, 2026, 2:50 p.m.