tgrep doesn't save tokens. Three lines in AGENTS.md do.
Tokens are getting more expensive. Providers still subsidise them, but the real cost of inference does not disappear — at some point it gets passed on. Anyone building agents into their work today is budgeting against a price that will be different tomorrow.
There is a second reason to be frugal, and it takes effect immediately: the fuller the context window, the worse the model works. Not once it overflows — well before. Anthropic calls it “context rot”. An agent that has loaded itself up with 400,000 tokens of search results is not just expensive, it is also worse.
Plenty is being offered accordingly. “Saves tokens” is printed on every second tool. We picked one of them, installed it and measured it, instead of copying the numbers from the announcement.
The result up front: the tool delivers what it promises — just not what it says on the box. And in the end we did not take it.
The map: three axes that cover everything
Look at the field and you find dozens of approaches that seem unrelated. They sort onto three axes, and that division helps place every new tool.
Axis 1 — put in less. What never enters the context costs nothing. This is where compression methods like Caveman belong, dropping grammar and keeping facts. It is also where deferred tool definitions belong — loaded on demand instead of carried permanently.
Axis 2 — keep less. What is in there need not stay forever. Summarising the transcript so far, notes in files instead of in the conversation, a memory that individual entries are pulled from.
Axis 3 — process elsewhere. The cheapest tokens are the ones another process spends. RAG systems, sub-agents with their own window, tools that compute rather than return raw data.

The fourth option is deliberately missing from that list: simply making the window bigger. That helps against the crash, not against the cost and not against the declining quality.
What a session actually costs
Before we get to tools, one number for scale. We asked an agent how many files sit in a particular folder. One question, one number as the answer, two turns.
Cost: 59,125 tokens and 0.32 US dollars. Four of those tokens were the question. The rest is system prompt, behavioural rules and tool descriptions — the base load, before anything happens at all.
Against that: a single careless search in a mid-size repository returns 116,121 tokens. Twice the entire base load, in one tool call.

The bottom bar is the path nobody takes on purpose: simply reading every matching file in full. 1,524,654 tokens. It is there because it marks the ceiling — and because an agent without a clear instruction drifts in that direction.
Behind the scenes: what we run
Before adopting a new tool, what is already running belongs on the table. The figures below are counted, not estimated.

Before the window. Our agents have access to 151 tool definitions across five MCP servers — studio API, Power Automate, browser control, memory, window selection. If every description were carried permanently, the window would be full before the first question arrives. Instead only the names are known; the description loads when the tool is actually needed. On top of that: skills, of which only one description line stays in context permanently, a read cache that never reads the same file twice, and a guard in front of every command that catches sprawling output.
Inside the window. When it fills up, the transcript so far gets summarised. The risk is that the summary discards the wrong thing, so three rules sit in front of it dictating what must survive: code changes, decisions, file paths. Plus 100 checkpoints from previous sessions and a memory of 24 individual files, of which only the index line loads at startup.
Beside the window. OpenViking as our own RAG on our own hardware. A knowledge wiki in the studio that stores the key findings of each article as a card — next time the card comes back, not the article. And sub-agents that search twenty files in their own window and report back three sentences.
One spot is open: searching the tree. There we search the way everyone does — across everything. That is exactly where tgrep would come in.
The newcomer: tgrep
It is microsoft/tgrep, published under the MIT licence and now shipping inside the GitHub Copilot CLI. At its core it is ripgrep with a pre-built index and a server in front.
Instead of reading every byte on every search, tgrep splits all text files once into trigrams — runs of three characters — and builds a sorted lookup table.
How a trigram works
You slide a three-character window across the text: “Hallo” yields Hal, all, llo. A text of n characters has n − 2 of them. That is the whole rule — no words, no grammar, no meaning.
The index reverses the direction. Instead of “file → contents” it stores “trigram → which files contain it”. Like the index at the back of a textbook, except the entries are not words but every three-character run that occurs.
Searching then means intersecting:

The sentence everything turns on: the intersection can never be larger than the shortest individual list. A single rare trigram in the pattern therefore discards nearly the whole tree. You need no statistics, just secondary-school set theory.
The guarantee holds in one direction only: anything containing urlaub necessarily also contains url, rla, lau and aub. So no match is ever lost. The reverse does not hold — a file can contain all four runs without containing the word. That is why the regex engine still reads the candidates at the end and decides. The index decides nothing, it only discards — and never anything correct. Which is precisely why tgrep returns character for character the same result as grep.
The limit belongs here too: a pattern needs at least three consecutive fixed characters. Search for .* or a.c and the index has nothing to intersect.
The architecture

The key question for understanding it is what actually runs permanently. The answer: exactly one layer. The server keeps the index memory-mapped and watches for file changes. The client the agent calls lives for milliseconds — it hands the query on over TCP and writes the answer to stdout. The regex engine does not run permanently but per query, and only across the candidate files.
The obvious assumption, and why it is wrong
The argument sounds compelling: an index saves context because fewer files get read. It still is not true.
You have to ask the same search two separate questions.

First: how long does the search take? This is where the index works. 311.6 milliseconds with ripgrep, 23.8 with tgrep — thirteen times faster.
Second: how many tokens come back? Here it does nothing at all. 116,121 with ripgrep, 116,121 with tgrep. The same number, because it is the same answer to the same question, character for character. That is why the top two bars on the right are exactly the same length.
The answer only gets shorter through a different question. -l instead of -n — file names only instead of every hit — takes the same search from 116,121 down to 9,076 tokens. And that works with ripgrep just as well, with nothing to install.
Three letters, three answer lengths
Because what follows keeps referring to “the rule”, here is concretely what that means. Nothing about ripgrep itself changes. It is the same command with a different letter:
rg -n "registerSingleton" # every hit: file, line number, line content
rg -c "registerSingleton" # only: which file, how many matches
rg -l "registerSingleton" # only: which files at all
Three commands, the same pattern, the same truth — but wildly different answer lengths. For the question “how many files contain this?” the middle line is entirely sufficient.
Without an instruction, an agent still reaches for the top one most of the time. It is the most verbose output, it is the obvious default, and it is what appears in the overwhelming majority of examples the model learned from.
The rule is therefore not a parameter but a sentence in the system prompt — it tells the agent which of the three lines to start with, and that it should take the top one only when it genuinely needs the line content. The agent still picks the parameter.
In Claude Code these three variants are called output_mode: "content", "count" and "files_with_matches". The same thing in different packaging.
Hands-on: the measurement
Everything that follows was measured with tgrep 1.0.8 and ripgrep 15.2.0, both as release binaries, on Windows 11. Timings are medians of five runs; the first run is discarded because it warms the file cache.
Setting up
# Windows: binary from the releases, otherwise Homebrew or cargo
tgrep index . # build the index
tgrep serve . # start the server, watches for changes
tgrep -- "pattern" . # search
Useful up front, because it settles the order of magnitude:
tgrep count-files .
# 385 text files (167 binary skipped, 0 too large, 0 errors) in 11ms
tgrep count-files . --no-ignore
# 65015 text files (1510 binary skipped, 0 too large, 0 errors) in 55ms
Two numbers for the same directory, a factor of 169 apart. The difference is solely whether .gitignore is honoured. Anyone quoting repository sizes should say which of the two they mean.
What the index costs
| Tree | Files | Build time | Index on disk |
|---|---|---|---|
| our repo (git space) | 706 | 236 ms | 12 MB |
our repo with node_modules | 66,792 | 6.6 s | 468 MB |
| vscode | 18,559 | 13.2 s | 187 MB |
| Linux kernel | 96,009 | 12.5 s | 1.1 GB |
Roughly a fifth of the source size as index. That appears in no announcement and belongs in every decision.
When it pays off
Two things decide it: how large the tree is and how many files the pattern hits. Read as a matrix:

The sweet spot is bottom left: large tree, rare pattern. There a handful of candidates survive out of 96,009 files, and that is exactly where the factor of 19 comes from.
The right-hand column shows the opposite. With static in the kernel nearly every second file matches — there is nothing to sort out, and managing the index costs more than it saves. The index helps exactly as far as it can discard. Both ends get more pronounced with tree size: the gain on the left grows from 2.1 to 19.4, the loss on the right from 1.0 to 0.3.
Below roughly a thousand files the effort does not pay off at all. That line matters again later.
In detail for the kernel:
| Pattern | Matching files | rg -n | tgrep -n | Factor |
|---|---|---|---|---|
devm_kzalloc | 5,759 | 1,377.8 ms | 70.9 ms | 19.4× |
EXPORT_SYMBOL_GPL | 3,664 | 1,239.4 ms | 139.6 ms | 8.9× |
static | 43,443 | 1,501.9 ms | 5,825.0 ms | 0.3× |
Checking the number from the README
The project documentation quotes 3,280 ms for ripgrep against 94 ms for tgrep on the Linux kernel under Windows, a factor of 34.8. On the same tree we measure 1,377.8 against 70.9 ms, a factor of 19.4.
The direction holds, the factor is smaller for us. Not because tgrep is slower — 70.9 against their 94 ms — but because our ripgrep runs more than twice as fast as theirs. Anyone quoting such factors is always also quoting the machine the counterpart ran on.
The actual test: 24 agent runs
The measurements above compare tools. The question that matters is a different one: does the same task cost less when the agent has a faster search?
Setup: the same task in the vscode repository — “In how many files does registerSingleton appear? Name the number and the three files with the most occurrences.” Model Sonnet, eight runs per variant, counting real tokens from the billing data, not estimates.
Three variants matter here, not two. Otherwise all you measure is that an agent with a rule works more frugally than one without:
| Variant | Tool | Rule in the system prompt |
|---|---|---|
| no rule | built-in search (ripgrep) | — |
| with rule | built-in search (ripgrep) | “file names or counts first, line contents later” |
| with rule + tgrep | tgrep via the shell | the same, plus how to call tgrep |
| Variant | Median | Mean | Spread | Turns |
|---|---|---|---|---|
| no rule | 322,108 | 339,528 | 16 % | 7 |
| with rule | 101,419 | 192,928 | 72 % | 2.5 |
| with rule + tgrep | 162,431 | 162,833 | 16 % | 3.5 |
All 24 runs produced the correct answer.
Three findings sit in there, and they appear to contradict each other:
The rule works harder than the tool. Without any new software the median drops from 322,108 to 101,419 — a factor of 3.2. That is the single largest effect in the entire series, and it comes purely from the agent using -c instead of -n.
But the rule alone is a gamble. The eight runs split into two groups: four at 79,009 tokens in two turns, four between 124,000 and 420,718 in up to nine turns. Sometimes the agent finds the short path immediately, sometimes it wanders.
The tool makes it predictable. All eight runs with tgrep land between 136,648 and 191,035. On the mean it is the cheapest variant; on the median it is not.
Compare only medians and you crown the rule. Compare only means and you crown tgrep. Both are here because both are true. In short:
The rule lowers the cost. The tool makes it predictable.
Two traps we walked into ourselves
A dormant index goes stale within minutes
Our first measurement run had no server. The result: tgrep --no-index found 18 matching lines, tgrep with the index only 16. At that point the index was around ten minutes old and did not know about two files created since.
No error message, no warning. Just a wrong answer.
With tgrep serve running and the file watcher active, all paths agreed. The server is not an accelerator, it is a precondition for correctness.
--no-ignore switches the index off
Our first series on the large tree suggested the index does nothing at all: identical times with and without it, both around 4,400 ms. The explanation is in the stats output:
tgrep --no-ignore --stats -l -- "wk_resturlaub" .
# Brute-force search completed in 4457.6ms (65016 files): 18 matches
tgrep --stats -l -- "wk_resturlaub" .
# 18 matches (18 matched lines) in 4.2ms (via server)
The same index, the same search, a factor of roughly a thousand. With --no-ignore, tgrep falls back to a full scan. Our first series had measured the index by switching it off — it is discarded.
The takeaway: an index covers exactly the space it was built for. Change the visibility rules while searching and you get no error, just a slow answer.
--stats is therefore the first command to reach for when a search runs slower than expected. Either it says “via server” — or it does not.
What you take away

Step 1: the rule
The most effective thing costs nothing and needs no installation. Three sentences in the agent’s instruction file:
## Searching the repository
When searching, first retrieve file names or match counts only
(output_mode "files_with_matches" or "count").
Pull line contents only when you actually need them.
Never read a whole file when a count will do.
The wording matters: this is an order, not a ban. Not “file names only, always” — at some point the agent needs the lines. The rule says where it starts, not what it may never do. A ban leads to it asking three times instead of reading properly once.
Where the file belongs. The cross-tool standard is called AGENTS.md, sits in the repository root and is read by more than thirty agents — Codex, Copilot, Cursor, Gemini CLI, Aider, Zed, Windsurf. The format is stewarded by the Agentic AI Foundation at the Linux Foundation. It is plain Markdown with no required fields.
Claude Code reads CLAUDE.md natively and by now AGENTS.md as well. To keep a single source, put @AGENTS.md as the first line of CLAUDE.md — that pulls the shared file in. A version for all projects lives at ~/.claude/CLAUDE.md.
Cursor additionally knows .cursor/rules/, Copilot additionally .github/copilot-instructions.md. You only need the tool-specific files for things that genuinely concern that one tool.
Incidentally, tgrep ships an AGENTS.md of its own, and it contains exactly this strategy: for broad searches use -l first, because the output is smaller. So Microsoft knows what matters, and writes it into the file almost nobody reads.
Step 2: add tgrep
Worth it from roughly ten thousand files. Below that the gain is small while the index is there regardless.
One thing is mandatory: the server has to run. Cleanest via a hook at session start that checks whether tgrep serve is up and starts it otherwise. If it is not running, the agent silently gets stale answers.
And on how to teach it to the agent: a rule is the right form for this, not a skill. A skill only loads when the agent considers it relevant — a search strategy has to apply on every turn, including when it does not notice it is searching. A skill is the right form only for the surrounding work: starting the server, checking index freshness, rebuilding a stale index.
Step 3: the rest
When the need arises: defer tool descriptions instead of carrying them all, sub-agents with their own window per subtask, an external memory that returns cards rather than full texts.
Limits
A few things this measurement does not show.
What was measured is a counting task. Whether the finding also holds for “find and change this” or “explain this flow to me” is open. For tasks that need a lot of context, the picture could look different.
All agent runs used one model. A model that plans better would make the with-rule variant steadier and might well consume tgrep’s advantage.
The search measurements ran on one machine, Windows 11 with twelve threads. On faster disks the index advantage shrinks — the comparison with the README already shows that.
And the most important blind spot: Dataverse development in large environments. Our trees are content repos with many small files. An unpacked solution export looks different — few, but very large XML and JSON files, plus generated early-bound classes and several environments side by side. We did not measure that. Anyone working there should not transfer the numbers below unchecked.
Conclusion: for our stack, we are not taking tgrep
That was not the plan. It was installed, it was measured, and the graphic for the stack was already finished.
Then we counted our own repos. The largest holds 389 text files. After that come 204, 163, 136. Not one reaches the order of magnitude at which an index takes effect.
In our own matrix that is the top row: a factor of 2.1. Twenty-five milliseconds become twelve. For that, build an index occupying a fifth of the code size? And run a server process that has to stay up for the answers to be correct at all? That is out of all proportion.

What we are adopting instead is the rule. Three sentences in a text file. In our measurement it achieved more than the tool: a task that previously cost 322,108 tokens costs 101,419 afterwards. A third. No installation, no running process, effective from the next start.
What this refusal does not cover. It concerns our content stack: blog, studio, reels and the agent pipeline behind them. Anyone doing Dataverse development in substantial customer environments faces a different tree, and the answer may well come out differently there. That is not a politeness clause but the honest scope of our measurement — we checked exactly twenty small repos and not a single large environment.
And we now know when we will fetch tgrep. As soon as a tree goes five digits. On vscode it was thirteen times faster, on the Linux kernel nineteen. That day may come — and the instructions for it are in this article.
So the one thing to take away is not the download. It is the three sentences.
See also
- A memory that stays: OpenViking as a personal RAG — the third axis at our place in detail
- Your flow is the tool — why agent tooling belongs where the logic already lives
Sources
- microsoft/tgrep — repository, README and the bundled
AGENTS.md, MIT licence - BurntSushi/ripgrep — the counterpart in every measurement
- Our complete measurement log with all raw values, the runner and the discarded series lives in the repo at
tools/token-messung/messprotokoll.md