Retrieval doesn't have to mean vectors. Coding agents proved it: they grep, read what they found, and grep again. StackGrep Engine gives your LLM the same loop over your documents, with no embedding model, chunking strategy or re-indexing job to maintain.
# your LLM's retrieval step, as a tool call
search_collection({ collection: "kb", query: "ERR_TLS_CERT_ALTNAME_INVALID" })
# or a pattern, when the shape matters more than the words
search_collection({ collection: "kb", query: "v[0-9]+\\.[0-9]+ (removed|deprecated)", regex: true })search_collection(query: "ERR_TLS_CERT_ALTNAME_INVALID")search_collection(query: "v4\.[0-9]+ deprecated", regex: true)get_document(id: "runbooks/tls.md")For lookups by name, ID, code or phrase it's better: an exact match is either there or not, and the agent can check it. For “find me something like this” questions, vectors still help, and the two work well together.
Let the model do it: an LLM is good at trying “cancel”, then “cancellation”, then “terminate”. Each search costs milliseconds, so a few in a row is still fast.
A pushed document is durable when the API answers and searchable within seconds (or before the answer, with wait=true). A synced bucket is checked every few minutes and on demand.
StackGrep Engine is onboarding teams from the waitlist. Bring your documents or your bucket.