Home / Blog

Claude Code greps. Your agent should too.

2026-10-04 · StackGrep

Ask Claude Code where a function is defined and watch what it does. It doesn't consult an embedding index. It runs a search for the name, reads the lines that come back, opens a file, and searches again. Cursor, Cline, Codex and Devin work the same way. A study of coding-agent source code found grep- and find-style tools in all eight agents it looked at (Inside the Scaffold).

That wasn't the obvious choice. A couple of years ago, retrieval for LLMs meant chunking documents, embedding them and running a nearest-neighbour search. Coding agents tried that and moved away from it.

Why coding agents chose grep

Boris Cherny, who created Claude Code, has described early versions using RAG with a local vector database, and the team finding that letting the agent search for itself worked better: simpler, with no index to go stale and nothing to leak. Cline wrote up the same decision. The reasons are worth spelling out, because they aren't specific to code:

  • Exact beats similar. A function name, an error code, an invoice number or a customer's name is either in the text or it isn't. Similarity search returns things that are close, and close is wrong when you're looking for INV-1042.
  • Results can be checked. A grep hit is a line of real text the agent can quote. The agent knows why it matched, and so do you.
  • Nothing goes stale. There's no embedding to recompute when a file changes.
  • The model does the clever part. LLMs are good at trying cancel, then cancellation, then terminate, and at narrowing with a pattern. Search is a tool in a loop, not a one-shot retrieval step.

Research is catching up. Is Grep All You Need? compares agent harnesses on a long-memory benchmark and finds grep-style search matching or beating vector retrieval. LlamaIndex weighed the trade-offs, and Doug Turnbull calls it agentic search's grep moment.

The objection: grep doesn't scale

The best argument against this came from Milvus: grep-only retrieval burns too many tokens. They have a point. Plain grep reads every byte on every search, and broad matches flood the context window. That's fine for a repository on a laptop. It's not fine for the data most agents outside coding need:

  • two million support tickets in a database export,
  • every contract the company has signed, as text,
  • 50 GB of logs and JSON sitting in an S3 bucket,
  • a product catalogue, a wiki, a decade of email.

None of that is on the agent's machine, and none of it can be grepped in milliseconds by reading it all. But the problem is reading everything, not grep itself.

Index the text, keep the grep

Code search engines solved this long ago. Google Code Search, and now Cursor's regex index for agent tools, keep an index of short byte sequences (n-grams). A search pulls the literal pieces out of the pattern, uses the index to find the few documents that contain all of them, and runs the real regex only on those. The answers are exactly what grep would return. Only the speed changes, from reading everything to reading almost nothing.

The token problem has a fix too: an agent rarely needs whole files. It needs the matching line, the document's id, and sometimes just a number. “How many tickets mention a chargeback?” should cost one call and one integer, not a context window full of tickets.

What this looks like over your data

This is what we built StackGrep Engine for. You push documents as JSON or point us at an S3 prefix, and your agent gets four tools over MCP or REST:

list_collections()
search_collection({ collection: "tickets", query: "chargeback", filter: "plan:enterprise" })
search_collection({ collection: "tickets", query: "INV-10[0-9]{2}", regex: true })
count_collection({ collection: "tickets", query: "chargeback" })   // exact, every match
get_document({ collection: "tickets", id: "t-88120" })            // word for word

Searches answer in milliseconds over hundreds of thousands of documents. Hits come back as the matching line with the document's id and metadata. Counts are exact. The index lives in object storage, so data you aren't searching costs storage and nothing else.

When you still want vectors

Embeddings are good at “find tickets like this one” when the words differ. If your agent's questions are about meaning rather than wording, keep a vector index for them. Most lookups an agent makes are names, IDs, phrases, codes and patterns, though, and for those it should do what coding agents do: grep, read and grep again, over all of your data, not just the files on one laptop.

See StackGrep Engine, read the quickstart, or try the same engine on the source of npm's top packages, no sign-up needed.

Give your agents exact search

StackGrep Engine is onboarding teams from the waitlist. Bring your documents or your bucket.

Join the waitlist Read the docs