Home / Engine / Use cases

Grep an S3 bucket: search the contents of every file

S3 can list your files but can't look inside them. Point StackGrep at a prefix: we index every text object, keep up as files change, and answer searches over all of it in milliseconds. Your data stays in your bucket; we only read it.

Join the waitlist See the docs

Why the usual way falls short

  • Downloading and grepping millions of objects takes hours and costs egress
  • Athena and S3 Select need a schema and scan everything, every time
  • Copying the data into a search cluster means a second copy to keep in sync

How it looks

PUT /api/collections/logs/source
{"bucket": "acme-logs", "prefix": "app/2026/", "region": "us-east-1"}

GET /api/collections/logs/search?q=OutOfMemoryError&filter=
→ every object that mentions it, with the line

What your agent asks, and what it calls

“Which files mention customer 88213?”search_collection(query: "88213")
“How many exports include a column named ssn?”count_collection(query: "\"ssn\"")
“Find every config that sets max_connections above 500”search_collection(query: "max_connections\s*=\s*([5-9][0-9]{2}|[0-9]{4,})", regex: true)

Questions

What access do you need?

Read-only access to the one prefix, granted by a bucket policy we show you, plus a verify file you write there so nobody can point us at a bucket that isn't theirs. Remove the policy and we stop.

Which files get indexed?

Every text object under the prefix, its key as the document id. Gzipped files are unzipped; binary files and objects over 4 MB are skipped.

Does my data leave my bucket?

We read it to build a compact index, kept in our object storage and cached where searches run. Delete the source and the collection goes with it.

More use cases

Give your agents exact search

StackGrep Engine is onboarding teams from the waitlist. Bring your documents or your bucket.

Join the waitlist Read the docs