Triple

T7906520
Position Surface form Disambiguated ID Type / Status
Subject Kolmogorov complexity E183589 entity
Predicate relatedTo P37 FINISHED
Object minimum description length principle
The minimum description length principle is a formal method in statistics and machine learning that selects the best explanation for data as the one that yields the shortest overall description of both the model and the data it encodes.
E700157 NE FINISHED

How this triple was built (4 steps)

Every LLM step that produced this triple, in pipeline order — named-entity classification, the disambiguation choices (the exact options shown, with the pick highlighted), and the generated description. The batch + timestamp of each is in the Provenance table below.

NER Named-entity recognition gpt-5-mini
Instruction
Given a phrase, classify it is english named entity (e.g., persons, organizations, works of art) in Latin script, or not (e.g., literals, dates, URLs, verbose phrases). For disambiguation, the statement where the phrase occurs as object is also given. Please return a JSON object with `phrase` (string, the phrase being analyzed) and `is_ne` (boolean, indicating whether the phrase is a Named Entity).
Input
Phrase: minimum description length principle | Statement: [Kolmogorov complexity, relatedTo, minimum description length principle]
NED1 Entity disambiguation (via context triple) gpt-5-mini-2025-08-07
Target entity: minimum description length principle
Context triple: [Kolmogorov complexity, relatedTo, minimum description length principle]
  • A. Kolmogorov complexity
    Kolmogorov complexity is a measure of the amount of information in an object, defined as the length of the shortest computer program that can produce it.
  • B. Bayesian Occam factor
    The Bayesian Occam factor is a term in Bayesian model comparison that automatically penalizes overly complex models by integrating over their larger parameter spaces, thereby implementing Occam’s razor in probabilistic inference.
  • C. Kullback–Leibler divergence
    Kullback–Leibler divergence is a fundamental information-theoretic measure that quantifies how one probability distribution differs from a reference distribution.
  • D. Fano inequality
    Fano inequality is a fundamental result in information theory that provides a lower bound on the probability of classification or decoding error in terms of conditional entropy.
  • E. Shannon–Khinchin axioms
    The Shannon–Khinchin axioms are a set of fundamental conditions that uniquely characterize Shannon entropy as the standard measure of information and uncertainty in probability theory and information theory.
  • F. None of above. chosen
  • G. Unsure - the case is ambiguous/there is not enough information to decide.
NEDg Description generation gpt-5.1
Instruction
Generate a one-sentence description of the target entity. 
You are given a context triple in the form (subject, predicate, object), where the object is the target entity. 
# Instructions
Use the triple to infer relevant information about the entity. Describe the entity based on what is most defining, well-known. 
Avoid repeating the information from the triple, unless really essential.
# Response Format
Return only the sentence: "Description: [one-sentence description of the target entity]"
Input
Entity: minimum description length principle
Triple: [Kolmogorov complexity, relatedTo, minimum description length principle]
Generated description
The minimum description length principle is a formal method in statistics and machine learning that selects the best explanation for data as the one that yields the shortest overall description of both the model and the data it encodes.
NED2 Entity disambiguation (via description) gpt-5-mini-2025-08-07
Target entity: minimum description length principle
Target entity description: The minimum description length principle is a formal method in statistics and machine learning that selects the best explanation for data as the one that yields the shortest overall description of both the model and the data it encodes.
  • A. Kolmogorov complexity
    Kolmogorov complexity is a measure of the amount of information in an object, defined as the length of the shortest computer program that can produce it.
  • B. Bayesian Occam factor
    The Bayesian Occam factor is a term in Bayesian model comparison that automatically penalizes overly complex models by integrating over their larger parameter spaces, thereby implementing Occam’s razor in probabilistic inference.
  • C. Kullback–Leibler divergence
    Kullback–Leibler divergence is a fundamental information-theoretic measure that quantifies how one probability distribution differs from a reference distribution.
  • D. Fano inequality
    Fano inequality is a fundamental result in information theory that provides a lower bound on the probability of classification or decoding error in terms of conditional entropy.
  • E. Shannon–Khinchin axioms
    The Shannon–Khinchin axioms are a set of fundamental conditions that uniquely characterize Shannon entropy as the standard measure of information and uncertainty in probability theory and information theory.
  • F. None of above. chosen

Provenance (5 batches)

The batch behind each pipeline step, in order, with when it ran. Timestamps are batch-level — stages were processed in waves, so the object chain (NER → NED1 → NEDg → NED2) reads in order, but predicate / elicitation batches can sit in a different wave.

Step Stage Batch ID Status When
creating Elicitation batch_69ca828dec0c81908b8f55a4dbbb53ff completed March 30, 2026, 2:02 p.m.
NER Named-entity recognition batch_69cb3a5871b8819087ad69c116c40091 completed March 31, 2026, 3:07 a.m.
NED1 Entity disambiguation (via context triple) batch_69cb5bc9dfa88190aa5261bdf44823ab completed March 31, 2026, 5:29 a.m.
NEDg Description generation batch_69cb7633c5a0819089deb6e89d9acb8e completed March 31, 2026, 7:22 a.m.
NED2 Entity disambiguation (via description) batch_69cbb84dc86c8190893d67ce07c51aa0 completed March 31, 2026, 12:04 p.m.
Created at: March 30, 2026, 5:03 p.m.