Triple

T1535776
Position Surface form Disambiguated ID Type / Status
Subject Berry–Esseen theorem E32545 entity
Predicate typicalMetric P10673 FINISHED
Object Kolmogorov distance
Kolmogorov distance is a statistical metric that measures the maximum difference between two cumulative distribution functions, commonly used to quantify convergence in distribution and in goodness-of-fit tests.
E174592 NE FINISHED

How this triple was built (5 steps)

Every LLM step that produced this triple, in pipeline order — named-entity classification, the disambiguation choices (the exact options shown, with the pick highlighted), and the generated description. The batch + timestamp of each is in the Provenance table below.

NER Named-entity recognition gpt-5-mini
Instruction
Given a phrase, classify it is english named entity (e.g., persons, organizations, works of art) in Latin script, or not (e.g., literals, dates, URLs, verbose phrases). For disambiguation, the statement where the phrase occurs as object is also given. Please return a JSON object with `phrase` (string, the phrase being analyzed) and `is_ne` (boolean, indicating whether the phrase is a Named Entity).
Input
Phrase: Kolmogorov distance | Statement: [Berry–Esseen theorem, typicalMetric, Kolmogorov distance]
NED1 Entity disambiguation (via context triple) gpt-5-mini-2025-08-07
Target entity: Kolmogorov distance
Context triple: [Berry–Esseen theorem, typicalMetric, Kolmogorov distance]
  • A. Kullback–Leibler divergence
    Kullback–Leibler divergence is a fundamental information-theoretic measure that quantifies how one probability distribution differs from a reference distribution.
  • B. Berry–Esseen theorem
    The Berry–Esseen theorem is a quantitative refinement of the central limit theorem that provides explicit bounds on the rate of convergence of normalized sums of independent random variables to the normal distribution.
  • C. Rényi divergence
    Rényi divergence is a family of information-theoretic measures that generalize Kullback–Leibler divergence to quantify the dissimilarity between probability distributions, parameterized by an order α.
  • D. Cameron–Martin theorem
    The Cameron–Martin theorem is a fundamental result in probability theory and functional analysis that characterizes how Gaussian measures on infinite-dimensional spaces change under shifts by elements of a special Hilbert subspace (the Cameron–Martin space).
  • E. Rényi entropy
    Rényi entropy is a generalized measure of information and uncertainty that extends Shannon entropy by introducing a tunable order parameter to emphasize different aspects of a probability distribution.
  • F. None of above. chosen
  • G. Unsure - the case is ambiguous/there is not enough information to decide.
NEDg Description generation gpt-5.1
Instruction
Generate a one-sentence description of the target entity. 
You are given a context triple in the form (subject, predicate, object), where the object is the target entity. 
# Instructions
Use the triple to infer relevant information about the entity. Describe the entity based on what is most defining, well-known. 
Avoid repeating the information from the triple, unless really essential.
# Response Format
Return only the sentence: "Description: [one-sentence description of the target entity]"
Input
Entity: Kolmogorov distance
Triple: [Berry–Esseen theorem, typicalMetric, Kolmogorov distance]
Generated description
Kolmogorov distance is a statistical metric that measures the maximum difference between two cumulative distribution functions, commonly used to quantify convergence in distribution and in goodness-of-fit tests.
NED2 Entity disambiguation (via description) gpt-5-mini-2025-08-07
Target entity: Kolmogorov distance
Target entity description: Kolmogorov distance is a statistical metric that measures the maximum difference between two cumulative distribution functions, commonly used to quantify convergence in distribution and in goodness-of-fit tests.
  • A. Kullback–Leibler divergence
    Kullback–Leibler divergence is a fundamental information-theoretic measure that quantifies how one probability distribution differs from a reference distribution.
  • B. Berry–Esseen theorem
    The Berry–Esseen theorem is a quantitative refinement of the central limit theorem that provides explicit bounds on the rate of convergence of normalized sums of independent random variables to the normal distribution.
  • C. Rényi divergence
    Rényi divergence is a family of information-theoretic measures that generalize Kullback–Leibler divergence to quantify the dissimilarity between probability distributions, parameterized by an order α.
  • D. Cameron–Martin theorem
    The Cameron–Martin theorem is a fundamental result in probability theory and functional analysis that characterizes how Gaussian measures on infinite-dimensional spaces change under shifts by elements of a special Hilbert subspace (the Cameron–Martin space).
  • E. Rényi entropy
    Rényi entropy is a generalized measure of information and uncertainty that extends Shannon entropy by introducing a tunable order parameter to emphasize different aspects of a probability distribution.
  • F. None of above. chosen
PD Predicate disambiguation gpt-5-mini-2025-08-07
Target predicate: typicalMetric
Context triple: [Berry–Esseen theorem, typicalMetric, Kolmogorov distance]
  • A. usesMetric
    Indicates that one entity adopts, applies, or relies on a particular metric or measurement standard in its operation, evaluation, or description.
  • B. typicalIn
    Indicates that something commonly occurs, appears, or is found within a given context, category, or environment.
  • C. isMetric chosen
    Indicates that something satisfies the properties required to be considered a metric, such as defining distances that obey non-negativity, identity, symmetry, and the triangle inequality.
  • D. primaryMetric
    Indicates the main quantitative measure used to evaluate the performance, success, or impact of an entity or process relative to its goals.
  • E. secondaryMetric
    Indicates that one metric serves as an additional, supporting measure used alongside a primary metric for evaluation or analysis.
  • F. None of above.

Provenance (6 batches)

The batch behind each pipeline step, in order, with when it ran. Timestamps are batch-level — stages were processed in waves, so the object chain (NER → NED1 → NEDg → NED2) reads in order, but predicate / elicitation batches can sit in a different wave.

Step Stage Batch ID Status When
creating Elicitation batch_69a885ea86308190998f6bc14bb91f8e completed March 4, 2026, 7:20 p.m.
NER Named-entity recognition batch_69a915f323bc8190aa757142c225e0ae completed March 5, 2026, 5:34 a.m.
NED1 Entity disambiguation (via context triple) batch_69ad295da0988190b1dc171bcdfe4d71 completed March 8, 2026, 7:46 a.m.
NEDg Description generation batch_69ad29dc9fd08190b67527f0662c92dc completed March 8, 2026, 7:48 a.m.
NED2 Entity disambiguation (via description) batch_69ad2a4d71a88190b67ed21beebbb479 completed March 8, 2026, 7:50 a.m.
PD Predicate disambiguation batch_69a907b046448190be8ea4d7b20255f7 completed March 5, 2026, 4:33 a.m.
Created at: March 4, 2026, 7:26 p.m.