Triple

T8285672
Position Surface form Disambiguated ID Type / Status
Subject Avi Kivity E193780 entity
Predicate employer P7 FINISHED
Object ScyllaDB
ScyllaDB is a high-performance, distributed NoSQL database designed as a drop-in replacement for Apache Cassandra, optimized for low latency and high throughput.
E724173 NE FINISHED

How this triple was built (4 steps)

Every LLM step that produced this triple, in pipeline order — named-entity classification, the disambiguation choices (the exact options shown, with the pick highlighted), and the generated description. The batch + timestamp of each is in the Provenance table below.

NER Named-entity recognition gpt-5-mini
Instruction
Given a phrase, classify it is english named entity (e.g., persons, organizations, works of art) in Latin script, or not (e.g., literals, dates, URLs, verbose phrases). For disambiguation, the statement where the phrase occurs as object is also given. Please return a JSON object with `phrase` (string, the phrase being analyzed) and `is_ne` (boolean, indicating whether the phrase is a Named Entity).
Input
Phrase: ScyllaDB | Statement: [Avi Kivity, employer, ScyllaDB]
NED1 Entity disambiguation (via context triple) gpt-5-mini-2025-08-07
Target entity: ScyllaDB
Context triple: [Avi Kivity, employer, ScyllaDB]
  • A. CockroachDB
    CockroachDB is a distributed SQL database designed for horizontal scalability, strong consistency, and high fault tolerance across multiple nodes and regions.
  • B. RethinkDB
    RethinkDB is an open-source, distributed NoSQL database designed for real-time applications by pushing live updates to clients as data changes.
  • C. MariaDB
    MariaDB is an open-source relational database management system, forked from MySQL, known for its compatibility, performance, and community-driven development.
  • D. MongoDB Cloud
    MongoDB Cloud is a fully managed cloud database platform that provides scalable, secure, and globally distributed MongoDB database services along with tools for data management, monitoring, and analytics.
  • E. TokuDB
    TokuDB is a high-performance storage engine for MySQL and MariaDB designed for large-scale, write-intensive workloads using Fractal Tree indexing to improve compression and insertion speed.
  • F. None of above. chosen
  • G. Unsure - the case is ambiguous/there is not enough information to decide.
NEDg Description generation gpt-5.1
Instruction
Generate a one-sentence description of the target entity. 
You are given a context triple in the form (subject, predicate, object), where the object is the target entity. 
# Instructions
Use the triple to infer relevant information about the entity. Describe the entity based on what is most defining, well-known. 
Avoid repeating the information from the triple, unless really essential.
# Response Format
Return only the sentence: "Description: [one-sentence description of the target entity]"
Input
Entity: ScyllaDB
Triple: [Avi Kivity, employer, ScyllaDB]
Generated description
ScyllaDB is a high-performance, distributed NoSQL database designed as a drop-in replacement for Apache Cassandra, optimized for low latency and high throughput.
NED2 Entity disambiguation (via description) gpt-5-mini-2025-08-07
Target entity: ScyllaDB
Target entity description: ScyllaDB is a high-performance, distributed NoSQL database designed as a drop-in replacement for Apache Cassandra, optimized for low latency and high throughput.
  • A. CockroachDB
    CockroachDB is a distributed SQL database designed for horizontal scalability, strong consistency, and high fault tolerance across multiple nodes and regions.
  • B. RethinkDB
    RethinkDB is an open-source, distributed NoSQL database designed for real-time applications by pushing live updates to clients as data changes.
  • C. MariaDB
    MariaDB is an open-source relational database management system, forked from MySQL, known for its compatibility, performance, and community-driven development.
  • D. MongoDB Cloud
    MongoDB Cloud is a fully managed cloud database platform that provides scalable, secure, and globally distributed MongoDB database services along with tools for data management, monitoring, and analytics.
  • E. TokuDB
    TokuDB is a high-performance storage engine for MySQL and MariaDB designed for large-scale, write-intensive workloads using Fractal Tree indexing to improve compression and insertion speed.
  • F. None of above. chosen

Provenance (5 batches)

The batch behind each pipeline step, in order, with when it ran. Timestamps are batch-level — stages were processed in waves, so the object chain (NER → NED1 → NEDg → NED2) reads in order, but predicate / elicitation batches can sit in a different wave.

Step Stage Batch ID Status When
creating Elicitation batch_69ca82e32db481908b72f3804fa71152 completed March 30, 2026, 2:04 p.m.
NER Named-entity recognition batch_69cb7ad20ae481908179aba245c73fad completed March 31, 2026, 7:42 a.m.
NED1 Entity disambiguation (via context triple) batch_69cd688441908190b6b0a39dfb9d87ac completed April 1, 2026, 6:48 p.m.
NEDg Description generation batch_69cd6d55196881909cf5ec925792e09f completed April 1, 2026, 7:09 p.m.
NED2 Entity disambiguation (via description) batch_69cd7e2bdae08190adc51e904e85695e completed April 1, 2026, 8:21 p.m.
Created at: March 30, 2026, 5:52 p.m.