Skip to content
Architecture2026-09-30

The Knowledge Graph My Agents Stopped Asking

A knowledge layer that agents consult but never obey fails into a bad suggestion instead of a bad commit.

In March 2026 I built a knowledge layer on Cognee and Neo4j so every agent in my stack had access to the decisions I had worked through in ChatGPT. I made it advisory from day one: agents consulted it, and it never overruled a repo, a test, or a written rule. Six months later the numbers say that call was right, and they say something harder too. Searches peaked at 64,951 in March and fell to one in September. The only agent still using it reads plain text chunks and never touches the graph.

The problem it solved

Most of my strategy and design decisions happened in ChatGPT conversations. The agents I build with (Claude Code, Codex, and a self-hosted agent) saw none of it. Every session started without the reasoning behind the systems it was working on.

This was before knowledge bases were a standard feature of every model product. Moving that history into one shared, queryable store was the most direct fix available. I built the integration so it did not matter which agent I ran that day; each one reached the same store. Cognee, an open-source graph memory framework, handled extraction, embeddings, and the graph build out of the box, and everything ran on my own hardware.

The hard part was the shape of the data. A ChatGPT conversation is a tree, not a transcript: every edited message forks it. Flatten the tree and you manufacture history, with two answers that never coexisted sitting in sequence as if one replaced the other. How I kept every branch intact is the subject of my NODES 2026 talk on November 12.

Advisory by design

I kept access narrow because I did not trust the output. By the time a conversation reaches the graph, it has been cut into chunks, turned into vectors, and handed to a model that decides what the entities are. That is three lossy steps between what I said and what an agent reads back. Useful as a hint. Dangerous as a source of truth.

So the rule, written into my agents' operating instructions, is that the graph is advisory retrieval context only. It has no authority over repo files, tests, or source code. When they disagree, the repo wins.

Running it proved the rule twice. The store's dataset filter does not isolate datasets, so a scoped search can return results from other collections. Later the corpus went stale while agents kept reading it every day. Both times the authority order held and only the advice was wrong. An advisory layer fails into a bad suggestion. An authoritative one fails into a bad commit.

What it does today, measured

Everything below comes from logs and read-only counts taken on September 30, 2026, covering the previous 30 days.

CallerRetrievals in 30 days
My self-hosted agent (read-only chunk recall, no graph)7 calls on 4 of 30 days, 5 successful: 1.7% of its 417 tool calls
Codex desktop1 search, across 128 connections that only listed the tools
Claude Code0
Cognee's search API, any caller0
Month (2026)Cognee searches
March64,951
April7,139
May1,345
June1,409
July1,047
August0
September1

The graph holds 8,707 nodes and 15,279 relationships, exactly what it held in July.

My read: this was the right tool in March and is mostly unneeded now. The models caught up. Long context, built-in memory, and instructions that live in the repo now carry what this layer used to carry. The one path still in use is the simplest one: a read-only text lookup.

What I changed once the numbers were in

  1. Removed the agent writers. An agent plugin was still configured to write its memories into the shared store. It had written nothing in 30 days, but advisory should not depend on an agent choosing not to write. Its write paths are off. I also removed the store from a desktop client that loaded its tools 128 times in a month and used them once.

  2. Stopped the monthly re-ingest. It had failed twice, each file waiting 35 minutes for GPU memory that a resident model never released. Fixing it means scheduling GPU windows around my video renders, which is not worth it for a corpus this lightly used. It is off, and turning it back on is one command.

What's next

  • Keep the graph running through the talk, then archive it with its backups instead of paying to run it.

  • Keep the read-only recall path only while it earns its calls. Right now that is seven a month.

  • Move the pieces that live only on the host into a repo: the recall reader, the extraction prompt, and the ontology.

We are all learning how much of this layer the models will absorb. Measuring beat guessing.