Knowledge graphs
A knowledge graph stores facts as labelled links between real-world things, so software can follow the links to answer questions.
Picture a web of dots and arrows. Every dot is a thing, such as a person, a place or a book, and every arrow names how two things relate. RDF is a standard from W3C, a group that sets web rules. In RDF, each fact is a triple: a subject, a predicate (the link) and an object. RDF calls a whole collection of these facts a graph. Other systems use a different model called a property graph.
The modern use of the name comes from Google’s 2012 announcement of its Knowledge Graph. At launch it covered over 500 million objects linked by over 3.5 billion facts and relationships. Wikidata is an open example: automated bots help add its data, and anyone may reuse it, even commercially, without asking.
A query is a question you ask the graph. You write each blank, called a variable, with a ? in front. The engine returns one row for each set of values that fits, so one question can give several rows. Wikidata puts wd: in front of an item and wdt: in front of a property. A slash joins a chain of links, so “educated at” and “part of” fit in one step.
Microsoft’s GraphRAG builds such a graph automatically. A language model pulls names, and the links between them, out of documents. Then, ahead of time, it writes a short summary for each tightly linked group. For a broad question, each summary drafts part of the answer, and a final pass merges the parts.
For decades, search matched the words people typed, not the real things those words name.
Follow one question through Wikidata, a public knowledge graph.
- 1 · identifyEvery item has a single ID made of the letter Q plus digits, so Douglas Adams is Q42.
- 2 · stateEach fact is a triple of subject, predicate and object, such as Q42, educated at (P69), St John's College (Q691283).
- 3 · linkSt John's College has facts of its own, including that it is part of the University of Cambridge.
- 4 · querySPARQL, a language for asking questions of knowledge graphs, writes the question as triple patterns with blanks called variables, and the engine fills in values that make those triples appear in the graph.
- 5 · answerFollowing educated at and then part of returns the University of Cambridge.
The answer comes from following two edges, not from finding one sentence that states it.
| Who | What they ask | What it works with |
|---|---|---|
| Web search | “taj mahal” | Which Taj Mahal is meant, the monument or the musician |
| Wikidata user | “Which university did Douglas Adams attend?” | The educated at edge, then the part of edge |
| News analyst | “Which big themes run through all these articles?” | Short write-ups of each cluster of linked names |
- One ID per thing lets a search engine tell Taj Mahal the monument from Taj Mahal the musician.
- Graph query languages can find things connected through paths of any length, not only direct links.
- Each fact can carry its evidence, since Wikidata records sources alongside its statements.
- Knowledge graphs are usually incomplete: many true facts are simply missing.
- A missing edge does not prove that the relation is false in the real world.
- When a language model builds the graph, it can record relationships that the source text never stated outright.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperKnowledge Graphs, Hogan et al., ACM Computing Surveys · read 29 Sept 2026
- officialRDF 1.1 Concepts and Abstract Syntax, W3C · read 29 Sept 2026
- officialIntroducing the Knowledge Graph: things, not strings, Google · read 29 Sept 2026
- docsWikidata:Introduction, Wikidata · read 29 Sept 2026
- docsWikidata:SPARQL tutorial, Wikidata · read 29 Sept 2026
- paperFrom Local to Global: A Graph RAG Approach to Query-Focused Summarization, Edge et al., Microsoft · read 29 Sept 2026
- officialSt John's College (Q691283), Wikidata · read 29 Sept 2026