Home / Engine / Use cases

Exact counts for RAG and agents

Ask a RAG system “how many customers asked for a refund?” and it counts the ten passages it retrieved. StackGrep's count reads every candidate and returns the exact number of matching documents, with a sample of their ids, in one call.

Join the waitlist See the docs

Why the usual way falls short

  • Top-k retrieval can't see past k, so counts and “all of them” questions come out wrong
  • Asking the LLM to count retrieved chunks is slow, costly and still incomplete
  • A database COUNT needs the data in tables first

How it looks

count_collection({ collection: "tickets", query: "refund", filter: "month:2026-09" })
→ {"docs": 1284, "sample": ["t-88120", "t-88131", ...]}

What your agent asks, and what it calls

“How many September tickets asked for a refund?”count_collection(query: "refund", filter: "month:2026-09")
“How many contracts renew automatically?”count_collection(query: "automatically renew")
“How many configs still pin TLS 1.0?”count_collection(query: "TLSv1(\.0)?[^.0-9]", regex: true)

Questions

Is the count exact or estimated?

Exact for your collections: every candidate is checked, and only documents that really match (and pass your filter) are counted.

How fast is a count?

Usually milliseconds to a few hundred milliseconds, depending on how many documents contain the words you're counting.

Can I count by metadata?

Yes: add a filter like customer:globex or month:2026-09, and run one count per value you want to compare.

More use cases

Give your agents exact search

StackGrep Engine is onboarding teams from the waitlist. Bring your documents or your bucket.

Join the waitlist Read the docs