CLASSEVE
RouteResearch
Research note · measured Sep 8, 2026 · published Sep 8, 2026

Code-graph context vs. file-by-file reading for coding agents: a 955-job measurement

Across 955 jobs on 18 public repositories (13,954 files, 2.7 million lines, 11 languages), a file-by-file agent workflow opened 1 file and read 1,413 lines — 13,321 tokens — for a typical job. The same job answered through Context Zero Engine's code graph used 2,221 tokens in one MCP call: an 83% reduction, 10.1× fewer tokens pooled across the run.

Setup

Eighteen public repositories — 13,954 files, 2.7 million lines, 11 languages, each pinned by commit — and one repository-analysis job repeated 955 times: change a function, which requires seeing the function, the code it uses from other files, and the code that calls it.

Two conditions. In the file-reading condition, the agent investigated by opening files and searching text, the default workflow of current coding agents. In the code-graph condition, the same questions were put to Context Zero Engine over MCP, which serves computed answers — symbols, callers, dependencies, blast radius — from a local PostgreSQL code graph.

Results

MetricFile-by-file workflowCode-graph workflow
Files opened10 (served from the graph)
Lines read1,413165
Tool calls21
Tokens consumed13,3212,221

The typical job is one real job: the median task in Alamofire, the repository whose saving equals the run's median of 83.3%.

The reduction comes from answering questions instead of transferring source: a caller list is a few hundred tokens, while the files containing those callers are tens of thousands.

Why this matters

An agent's context window is its working memory. Tokens spent re-reading source to rediscover structure are tokens unavailable for the actual change. Persisting that structure in a queryable graph converts orientation from a per-session cost into a one-time index.

Scope.

  • Savings vary by repository, from 2.6x on Express to 16.5x on Django; on 101 of the 955 jobs, reading the files directly was cheaper.
  • The benchmark ships in the repository, so the run can be repeated on any codebase.