sys1grep — System 1 for your files; uses TypeSafe AI's Jev
Background and design notes (Japanese): Jevのキラーアプリ、「意味で探す grep」を作った on Zenn.
Renamed from semgrep (0.5.0). Two reasons. The name collided with the static-analysis tool Semgrep, a trademark. And the tool was never meant for Jev alone: it targets System 1 models in general, and I'd like it to work with any models TypeSafe AI may release after Jev under other names :). So from 0.5.0 this is sys1grep (System 1 + grep): npm package@uehaj/sys1grep(was@uehaj/semgrep), commandssys1grep/git sys1grep, environmentSYS1GREP_, config~/.config/sys1grep/.env. The oldSEMGREP_names and~/.config/semgrep/.envstill work for one minor release, each use printing a deprecation line; 1.0.0 drops them. Details: CHANGELOG. The Zenn article's URL keeps its old slug.
A grep that finds lines by what they mean, not by regular expressions. Matching is done by Jev, the System One model from TypeSafe AI. Jev does not generate text. It answers typed questions with probabilities, so for every line sys1grep asks "does this line match the meaning network failure?", gets a probability back, and applies a threshold.
./sys1grep -n -e "customer is angry or frustrated" tickets.txt
▶ Click the demo, or open uehaj.github.io/sys1grep, for the full demo on the landing page.
- Zero dependencies. One file, Node.js 20.16+ and
fetch. - Fast. 30 lines go into one request, requests run 8 at a time. A 210-line file finishes in under a second.
- Meanings combine with AND / OR / NOT.
- Language-agnostic. The meaning and the text can each be in any language. A Japanese meaning finds French, Russian, Chinese and Korean lines alike. No translation step, same speed, same cost.
Search across languages
The meaning and the text do not have to share a language. Jev compares concepts, not words, so one query finds matching lines in every language the file contains.
An English meaning finds Japanese lines. None of the hits contain "angry" or "frustrated", and two of them are in Japanese:
$ ./sys1grep -n -e "customer is angry or frustrated" tests/corpus.txt
14:ユーザー山田さんからの問い合わせ: 返金してほしい、商品が壊れていた
16:ユーザー佐藤さんからの問い合わせ: 注文した覚えのない請求が来ています。至急確認してください
18:I want my money back. The item arrived broken and customer service ignored me.
21:Your product ruined my weekend. Never buying from you again.
23:This is the third time I'm writing. Nobody has replied to my previous emails.
A Japanese meaning finds English lines, with the same confidence as the Japanese ones:
$ ./sys1grep -n -p -e "返金の要求" tests/corpus.txt
14:ユーザー山田さんからの問い合わせ: 返金してほしい、商品が壊れていた [0.97]
18:I want my money back. The item arrived broken and customer service ignored me. [0.95]
It is not a Japanese/English feature. tests/multi.txt holds refund requests and thank-you notes in French,
Russian, German, Spanish, Chinese and Korean. One Japanese meaning finds all six refund requests; a Russian
meaning does the same:
$ ./sys1grep -n -p -e "顧客が返金を求めている" tests/multi.txt
1:Je veux être remboursé, le produit est arrivé cassé. [0.98]
3:Я требую вернуть деньги, товар не работает. [0.97]
5:Ich möchte mein Geld zurück, das Gerät ist defekt. [0.97]
7:Quiero un reembolso, el paquete llegó vacío. [0.97]
9:我要求退款,商品坏了。 [0.97]
11:환불해 주세요. 제품이 고장났어요. [0.97]
$ ./sys1grep -n -p -e "клиент требует возврат денег" tests/multi.txt
1:Je veux être remboursé, le produit est arrivé cassé. [0.94]
3:Я требую вернуть деньги, товар не работает. [0.97]
5:Ich möchte mein Geld zurück, das Gerät ist defekt. [0.92]
7:Quiero un reembolso, el paquete llegó vacío. [0.89]
9:我要求退款,商品坏了。 [0.92]
11:환불해 주세요. 제품이 고장났어요. [0.89]
This makes sys1grep useful for mixed-language logs and ticket dumps, and for teams whose members query in different languages. One caveat: TypeSafe documents English as the most accurate language, and in our tests Japanese meanings wobble a little more near the threshold. When a query is borderline, phrasing the meaning in English is the safer choice.
How this differs from vector search
If all you want is "lines about X", embedding similarity gives you the same lines. sys1grep differs in what it judges: not how close a line is to a topic, but whether a proposition holds for that line. Jev reads the line and the question together (a cross-encoder shape), so who did what, negation, and "asked for" versus "already done" all change the answer. An embedding of the line is fixed before it ever sees your query, so it can only measure topical closeness.
All eight lines below are "about a refund" (the same contrast as tests/contrast.txt, translated to
English as tests/contrast.en.txt — see Search across languages above for
the case where the meaning and the text differ in language). Only two are a customer asking for one:
$ ./sys1grep -n -p -t 0 -e "customer is asking for a refund" tests/contrast.en.txt
1:I want a refund. The item was broken. [0.99]
2:The refund has been processed. Please check your account. [0.11]
3:Our refund policy is within 30 days of purchase. [0.09]
4:The manager denied the refund request yesterday [0.17]
5:I demand a full refund immediately [0.93]
6:Refunds are processed within 5 business days [0.07]
7:The support agent got angry and hung up the phone. [0.05]
8:The customer got angry and hung up the phone. [0.06]
Because each meaning yields an independent probability, **logical AND and NOT are plain boolean operations**, not a trick with set differences or "negative queries":
# about a refund, but NOT a customer asking for one → completed, policy, denied, timelines
$ ./sys1grep -n -e "about a refund" -v "the customer is asking for a refund" tests/contrast.en.txt
2:The refund has been processed. Please check your account.
3:Our refund policy is within 30 days of purchase.
4:The manager denied the refund request yesterday
6:Refunds are processed within 5 business days
angry AND it is the customer, not the staff
$ ./sys1grep -n -e "someone is angry" -a "the customer, not the staff, is the one acting" tests/contrast.en.txt
5:I demand a full refund immediately
8:The customer got angry and hung up the phone.
Line 7, The support agent got angry and hung up the phone., scores 0.05 on the second meaning and is
excluded, even though it differs from line 8 by one word (support agent vs. customer) — wording an
embedding would place right next to line 8's. (We could not reproduce a cosine-similarity number for
this pair: no embedding-model access was available in this environment, and no prior measurement script
exists in the repository or its history to rerun. The argument stands on the wording alone: an embedding
of line 7 has no way to see that "support agent" changes who the sentence is about.)
Two more practical consequences. The probabilities are calibrated, so one threshold (0.5) works across queries, where cosine scores need top-k or per-query tuning. And there is no index to build: sys1grep reads the files in front of you. The flip side is that every query pays for the whole corpus again, so for repeated queries over a large, fixed corpus a vector index is cheaper and faster.
Regex terms
-e/-a/-v '/pattern/flags' (first and last character /, JavaScript flags) is matched locally,
as a plain regex, with no request at all. It prefilters its AND term: only the lines it holds for
ever ask that term's meanings, so a cheap regex in front of a meaning cuts both the bill and the wait. A line is sent
only if some term's regexes all hold for it (a term with no regex holds for every line), so with
-e '/re/' -a A -e B a line without re is still sent, asked B only. A query of regex terms alone
sends nothing, even with --unit=sentence-by-jev: the rules alone then decide where wrapped lines join. Anything that isn't shaped like /…/flags is still a meaning, so
-e '/etc 以下のファイルを変更している' (no closing /) is unaffected; a meaning that really starts
and ends with / can be written with a leading space to dodge the regex reading.
$ ./sys1grep -e '/ERROR|FATAL/' app.log # no requests at all
$ ./sys1grep -e '/timeout/i' -a '顧客に影響が出ている' app.log # only lines with "timeout" go to Jev
A regex's named and numbered groups pass to the other meanings of the same AND term as
$, $1-$99, $&, $$ — ECMAScript's replacement-pattern syntax
(String.prototype.replace's GetSubstitution), with one deviation: $ naming no group is
an error rather than an empty string (so is a reference to a negated regex's group). A $n naming
no group stays literal, as in ECMAScript, so $100 以上の請求 is unaffected.
$ ./sys1grep -e '/(?<date>\d{4}-\d\d-\d\d) (?<time>\d\d:\d\d)/' \
-a '$<time> が深夜(0時〜5時)であり、$<date> が週末である' app.log
2026-09-19 03:12 ... → asks "03:12 が深夜(0時〜5時)であり、2026-09-19 が週末である"
Prefer $ and single quotes: $ survives double quotes in sh/bash/zsh; $1, $time
and ${time} don't (the shell expands them itself). -p prints 1.00/0.00 for a regex term.
On a terminal every match of a regex is in grep's match color (bold red), for the terms that held; a negated regex
is never colored. -o prints each match on a line of its own, as grep -o; a line only meanings matched prints
whole, since a meaning has no matching part.
$ ./sys1grep -o -n -e '/[A-Z]+-\d+/' -a 'the ticket is still open' notes.txt # the ticket ids, one per line
Sending less
Every line sent to Jev costs money and time, so the cheapest line is the one never sent. From the widest cut to the narrowest:
- Which files.
-rskips.git,node_modules, binary files, likely secrets, generated files (source maps,
git sys1grep searches
tracked files only. --include / --exclude (file-name globs) and --changed-within (30m, 7d, today,
this-week, a date) narrow them further. A meaning that restricts itself to a language or a time of change
narrows them by itself: see Scope from the meaning.
- Which lines. A regex term is matched locally, and only the lines it holds for are asked
-M/--max-columns characters of a line
(default 2000, 8000 with -z) are sent: it is still searched and judged, just not past that cutoff.
- How many times.
--dedupjudges one line per template: lines that
auto/always/never, off by default for
now: a hint says when --dedup=auto would pay.
- Check before paying.
--dry-runsends nothing and prints the settings the search would run with, the
~3178 input tokens, ~$0.000133; within about 10%). -i shows the
same totals on the terminal and sends only after y.
Before the bulk of requests, every target is sized (--max-filesize, default 10M); one over it is skipped
outright, like rg's own --max-filesize, named on stderr (-y does not affect it). A .gz is sized
decompressed (zlib stops at the limit, so an oversized one is never inflated in full). stdin is sized once it
is read (it is already in memory), and skipped the same way if it is over; a --cached or : target
is a blob rather than a file, sized from its content (already read in one git cat-file --batch call, not a
fresh git process per blob); -g's commits stay out of this, since each is already bounded by -M when it
is sent. The input about to be sent is priced too (--max-cost,
default 1 USD); over it, one question asks to continue on the terminal, -y answers it yes, and without a
terminal it is exit 2 (this applies with -q too — scripts pass -y).
With -r or git sys1grep, a term no regex narrows that is about to send more than 10,000 lines gets one
stderr line with the same totals, naming the term (sys1grep: sending 214,913 of 231,502 lines from 1,247 files
(~9.1M input tokens, ~$0.38); the term "…" has no regex to narrow it. Add -a '/RE/' to it, …). The search goes
on; -q silences the line, -y does not.
$ sys1grep --dry-run -r --include='*.log' --changed-within=today -e '/ERROR|FATAL/' -a 'a customer is affected' logs/
--verbose (or --dry-run) also shows where a setting that was not typed on the command line came from
(SYS1GREP_OPTS, an environment variable, ~/.config/sys1grep/.env, or a preset's default), so a result that
surprises you can be traced back to its source. The key's value never appears, only which option or variable
supplied it:
$ SYS1GREP_OPTS='--level strict' sys1grep --verbose -e "the API key is read from a file" .
sys1grep: endpoint api.typesafe.ai/v1/systemone (default), model jev-latest (default)
sys1grep: key: SYS1GREP_API_KEY (~/.config/sys1grep/.env)
sys1grep: SYS1GREP_OPTS: --level strict
sys1grep: options: --level strict (SYS1GREP_OPTS) = -t 0.7 -T 0.3, --chunk 30, -j 8, scope on, -M 2000, --max-filesize 10M, --max-cost 1
sys1grep: file ./a.py: 120 lines, 120 to send
…
Install
Two ways to use it: as a command-line tool (this section), or as a Claude Code skill
(see Use it from Claude Code below). The skill falls back to npx @uehaj/sys1grep,
so if you only use it through Claude Code you can skip the install here entirely and just set the API key.
Requires Node.js 20.16 or later. No other dependencies.
npm install -g @uehaj/sys1grep
sys1grep --help
To try it without installing, run it through npx (the first run downloads the package, later runs use the cache):
npx @uehaj/sys1grep -n -e "customer is angry or frustrated" tickets.txt
Then give it an API key from the TypeSafe console. Any one of these works:
export SYS1GREP_API_KEY=your-key # environment variable
echo 'SYS1GREP_API_KEY=your-key' > ~/.config/sys1grep/.env # per user (mkdir -p first)
Variables already in the environment win; otherwise ~/.config/sys1grep/.env fills them in. A .env in the current
directory is never read: it may belong to a repository you just cloned, and could send your key elsewhere through
SYS1GREP_URL. For per-project settings, load a file yourself: node --env-file=.env "$(command -v sys1grep)" ....
TYPESAFE_API_KEY is accepted too when SYS1GREP_API_KEY is not set.
Default options
SYS1GREP_OPTS holds options to apply on every call, read like the settings above. It is split on spaces and put in
front of the command line, so the command line wins: a later value counts, and --no-X turns a default flag off.
export SYS1GREP_OPTS='--level strict -j 8 -n'
sys1grep -e "payment failed" app.log # strict, 8 at once, line numbers
sys1grep --level loose --no-n -e "payment failed" app.log
A script calling sys1grep would pick these up too (grep dropped GREP_OPTIONS for that reason). Call it as
SYS1GREP_OPTS= sys1grep ... in scripts.
Other endpoints
The API is configured by exactly three settings: SYS1GREP_API_KEY (or TYPESAFE_API_KEY), SYS1GREP_URL and SYS1GREP_MODEL.
Any endpoint that speaks TypeSafe's POST /v1/systemone works. The key is sent to SYS1GREP_URL as is, so set the
two together. On the command line, --sys1-model=ID, --sys1-url=URL and --sys1-api-key=KEY override
the three. A key given this way shows up in ps and shell history, so prefer ~/.config/sys1grep/.env for it.
# OpenRouter
SYS1GREP_URL=https://openrouter.ai/api/v1/systemone SYS1GREP_API_KEY=sk-or-... sys1grep -e ...
Vercel AI Gateway
SYS1GREP_URL=https://ai-gateway.vercel.sh/typesafe/v1/systemone SYS1GREP_MODEL=typesafe-ai/jev SYS1GREP_API_KEY=vck_... sys1grep -e ...
A compatible server that needs no key: no Authorization header is sent
SYS1GREP_URL=http://localhost:8000/v1/systemone sys1grep -e ...
The summary line shows the cost the endpoint reports (usage.cost), or for TypeSafe itself an estimate at list
price marked ~.
From source: git clone https://github.com/uehaj/sys1grep.git && cd sys1grep && npm install -g .,
or run it in place with node sys1grep.mjs ....
Examples
All examples run against tests/corpus.txt, a 51-line mix of server logs,
support tickets in English and Japanese, source code, SQL and small talk. Where a Japanese line would
otherwise show up in the output below, this section instead uses tests/corpus.en.txt,
the same 51 lines with the Japanese ones translated to English — the cross-language behaviour itself is
shown once, in Search across languages above.
Find lines by a concept, in any language
$ ./sys1grep -n -e "customer is angry or frustrated" tests/corpus.en.txt
14:Inquiry from user Yamada: I want a refund, the item was broken
16:Inquiry from user Sato: I'm being charged for an order I never placed, it looks like fraud, please check urgently
18:I want my money back. The item arrived broken and customer service ignored me.
21:Your product ruined my weekend. Never buying from you again.
23:This is the third time I'm writing. Nobody has replied to my previous emails.
5 of 51 lines matched; 51 sent to Jev in 2 requests, 3054 input tokens, ~$0.000128
None of these lines contain the words "angry" or "frustrated".
Lines that answer a question (-Q)
-e asks whether a line states a meaning. -Q QUESTION (--question) finds lines that answer it instead,
so you can write the question as you would ask it:
$ echo "The job failed." | ./sys1grep -e "Did the job succeed?"
$ echo "The job failed." | ./sys1grep -Q "Did the job succeed?"
The job failed.
As a meaning, "Did the job succeed?" is not what the line says (0.08), so -e finds nothing. As a question,
the line answers it: no, it failed (0.83). A line that asks the question is close to it in meaning but
answers nothing, and for a yes/no question a line that denies it still answers it:
$ ./sys1grep -n -Q "whether the server is down" tests/intent.txt
5:The server is down.
6:The server is healthy and responding normally.
2 of 17 lines matched; 17 sent to Jev in 1 request, 1087 input tokens, ~$0.000046
Both the confirming line and the denying line match: each settles whether the server is down. Is the
server down? does not match, because asking is not answering; -e "asking whether the server is down"
would match it instead. -Q X is shorthand for -e "the line answers: X", so it combines with -a / -v
/ ! and OR's with other terms exactly like -e.
OR: two meanings, and see the probabilities with -p
$ ./sys1grep -n -p -e "customer is asking for a refund" -e "delivery address change request" tests/corpus.en.txt
14:Inquiry from user Yamada: I want a refund, the item was broken [0.99 0.01]
17:Inquiry from user Takahashi: I'd like to change the delivery address [0.01 0.99]
18:I want my money back. The item arrived broken and customer service ignored me. [0.96 0.01]
22:Can I change the delivery address for order #8821? [0.02 0.99]
4 of 51 lines matched; 51 sent to Jev in 2 requests, 4275 input tokens, ~$0.000180
The bracket shows one probability per meaning, in the order given. Use it to pick a threshold.
With --color (on by default in a terminal) the probabilities are colored against the thresholds:
green at or above -t, red below -T, yellow in between. Line numbers and file names use grep's colors.
!colored output: line numbers in green, probabilities in green or red
AND NOT: network errors, excluding retries
$ ./sys1grep -n -e "network or remote connection failure" -v "a retry is happening or was attempted" tests/corpus.txt
4:2026-09-19 08:02:30 ERROR connection reset by peer while calling payment-gateway
6:2026-09-19 08:02:35 ERROR timeout after 5000ms waiting for payment-gateway
9:2026-09-19 08:10:44 ERROR DNS lookup failed for api.example.com
11:2026-09-19 09:00:00 ERROR SSL handshake failed: certificate expired
13:network unreachable: no route to host 10.0.0.5
30:except ConnectionError as e:
31: logger.error("upstream unreachable: %s", e)
7 of 51 lines matched; 51 sent to Jev in 2 requests, 5112 input tokens
Line 5, retrying payment-gateway request (attempt 2/3), is a network failure but is dropped by -v.
Mixed: (finance AND negative) OR weather
$ ./sys1grep -n -e "about economy, finance or markets" -a "the news is negative or a decline" -e "about weather" tests/corpus.en.txt
36:Today's weather is sunny, high of 28 degrees
43:Stock prices fell 3% after the earnings report missed expectations.
48:Tomorrow's forecast is rain so I'll bring an umbrella
3 of 51 lines matched; 51 sent to Jev in 2 requests, 5499 input tokens, ~$0.000231
The central bank raised interest rates is about finance but not a decline, so it is out.
Strictness presets
$ ./sys1grep --level strict -n -e "a security risk or dangerous destructive operation" tests/corpus.en.txt
33:DROP TABLE sessions;
49:API keys must never be committed to the repository.
$ ./sys1grep --level loose -n -e "a security risk or dangerous destructive operation" tests/corpus.en.txt
11:2026-09-19 09:00:00 ERROR SSL handshake failed: certificate expired
16:Inquiry from user Sato: I'm being charged for an order I never placed, it looks like fraud, please check urgently
33:DROP TABLE sessions;
49:API keys must never be committed to the repository.
strict keeps only what the model is sure about. loose also pulls in the expired certificate and the suspicious-billing ticket.
Recursive search and file names only
$ ./sys1grep -r -n -e "customer is asking for a refund" tests/tickets/
tests/tickets/a.txt:7:Ticket #16: I want a refund, the item was broken.
tests/tickets/sub/b.txt:1:The customer wants a refund for the broken lamp.
$ ./sys1grep -rl -e "customer is asking for a refund" tests/tickets/
tests/tickets/a.txt
tests/tickets/sub/b.txt
-r walks directories in sorted order and skips .git, node_modules, .ssh, .aws, .gnupg, .kube,
.docker, binary files (a NUL byte in the first 8 KB, or a PDF; UTF-16 with a BOM is read as text), files
that usually hold secrets (.env, .netrc, .npmrc, .pypirc, .pgpass, .git-credentials, .pem,
.key, .p12, .pfx, .jks, .keystore, id_rsa and friends; names compared without case) and
generated files (.map, .min.js, *.min.css, package-lock.json, yarn.lock, pnpm-lock.yaml,
Cargo.lock, poetry.lock, composer.lock, Gemfile.lock, go.sum). **Every line that is searched is sent
to the TypeSafe API**, so point -r at a directory
you mean to scan. Inside a git repository, -r also skips what git ignores (.gitignore, .git/info/exclude, the
global excludes file), so build output and local files stay home; a tracked file is searched even if it matches.
A file named *.gz (a rotated log: app.log.1.gz) is read decompressed, as zgrep does, and printed by its
name on disk (app.log.1.gz:12:...); --include='*.gz' picks those, the binary sniff sees the decompressed
bytes, a gzipped secret (private.key.gz) is skipped like the plain one, and a corrupt .gz is an unreadable file.
Only gzip, only by the name.
A file named explicitly on the command line is always searched, even if it matches the skip list or is ignored;
so is a directory that git ignores, when you name it (sys1grep -r -e ... dist). -l prints each
matching file once, in the order matches are found, and works with or without -r. -c prints the number of
matching lines per file instead.
Scope from the meaning
A meaning that restricts its matches to some kind of file can only match in such files. With -r and
git sys1grep, each meaning first asks Jev, in one small request, a yes / no per candidate, and a yes at 0.6 or
more narrows the files before anything else is sent. The rest is never read or sent. Each scope is reported on
stderr with Jev's answer, so a wrong one is visible:
$ sys1grep -r -e 'Python でリトライ処理を書いている箇所' .
sys1grep: scope: .py .pyi *.pyw (from "Python files: 0.94")
sys1grep: scope: 12 of 340 files
$ sys1grep -r -e '昨日変えた箇所で認証を扱っている' src/
sys1grep: scope: modified since 2026-09-25 00:00 (from "what was changed yesterday: 0.91")
sys1grep: scope: 3 of 120 files
- Language or format: 26 candidates (Python, JavaScript, TypeScript, Go, Rust, Java, Kotlin, Ruby, PHP, C, C++,
- Time of change: 14 spans: the last minute, the last hour, today, yesterday, the last 1 / 2 / 3 / 7 / 30 days,
- Place: test code (
tests/ test/ __tests__/ spec/ e2e/,test_.py _test. .test. .spec. Test.java
.rs: Rust keeps unit tests inside the file), database migrations (migrat, Flyway's
V1__.sql), the README, the changelog (CHANGELOG CHANGES HISTORY NEWS), documentation (.md .rst .adoc
*.txt, docs/), source code (whatever is not documentation, so languages missing from the dictionary still
count) and logs (.log .log.N .out .err, logs/), by the path conventions of JS, Python, Go, Java, Ruby, Rust
and PHP. Several places are alternatives ("README か CHANGELOG に").
- git: in a repository, a time goes by commits: a committed file needs a commit at or after the start (the
origin/HEAD, main
or master), not pushed (@{upstream}..HEAD, else commits on no remote branch), and mine (user.email's
commits and the uncommitted files). Authors: the 30 with the most commits (git shortlog), each a candidate by
name and e-mail; a yes narrows to every file a commit of theirs touched. Outside a repository none of these is
asked, and a time goes by the mtime.
- Jev reads the whole meaning, in any language: "案A、B、Cで比較" is not about C files, and a date quoted in a
- Per term.
-e A -e Bstill searches B in the files A's scope leaves out; within an AND term, and across
-v, !) are not asked. The
meaning is sent unchanged.
- The question costs one small request per meaning, sent only when
-r/git sys1grepfound something to
-i's answer. --dry-run shows it as [scope]. As with --include, with -r a file named on
the command line and stdin are never narrowed, and git sys1grep's pathspecs are narrowed like the rest. --no-auto-scope turns it off (--auto-scope turns it back on).
As a git subcommand (git sys1grep)
npm install -g also installs git-sys1grep, so git runs it as git sys1grep. Like git grep, it searches only
the files git tracks (ignored files and build output are never sent), and FILE arguments are pathspecs relative
to the current directory.
$ cd tests && git sys1grep -l -e "customer is asking for a refund" fixture.txt tickets
fixture.txt
tickets/a.txt
tickets/sub/b.txt
Without FILE it searches every tracked file under the current directory. The -r skip list (.env*, keys, ...)
applies even to tracked files. --include, --exclude and --changed-within filter the tracked files, those named
by a pathspec included. --changed-within reads the working tree's modification times, not git history: right after
a clone or a checkout, every file it wrote counts as just changed. For help use git sys1grep -h: git takes --help itself and looks for a man page.
--cached, --untracked and a choose what is searched instead of the working tree, as git grep has
them (only one of the three at a time):
$ git sys1grep --cached -e "a retry is attempted" # staged, uncommitted changes included
$ git sys1grep --untracked -e "a retry is attempted" # tracked files plus untracked ones (.gitignore still applies)
$ git sys1grep -n -e "a retry is attempted" main v0.3.1 -- '*.py'
main:src/job.py:42: retry(job, times=3)
v0.3.1:src/job.py:40: retry(job)
A (a branch, tag, commit or @{u}) is any argument before -- that resolves as a revision; several may
be given, searched in the order given, each line prefixed with the name as typed (@{u}:path, not the branch it
resolves to). --changed-within needs the working tree: a blob (--cached, a ) has no mtime of its own.
A 's own pathspec takes a glob too, same as elsewhere in sys1grep (git diff-tree against the empty
tree, not git ls-tree's own literal/directory-prefix match); --include / --exclude (by name) still work.
An old can hold a secret an ordinary file once carried and was later removed from: the skip list
drops files by name only, not by what changed since.
Everything that is not something
./sys1grep -v "a timestamped server log line" mixed.txt # like grep -v
./sys1grep -e "source code or SQL" -v "SQL" src.txt # code, but not SQL
cat app.log | ./sys1grep -e "the deploy failed or was rolled back"
Records that span several lines (-z)
The unit of judgement is a line. That is right for logs and source, and wrong when one record spans
several lines. -z makes the unit a NUL-terminated record instead, exactly as in grep -z, so it pairs
with the tools that already emit records: git log -z, find -print0, xargs -0.
A proposition like "this commit changes user-visible behaviour" is true of a whole commit, not of any one line in it:
$ git log -z --format='%h %s %b' | ./sys1grep -z -n -e "the change alters user-visible behaviour" -v "documentation only"
7:21120e9 Revert "feat: ship the /sys1grep Claude Code skill" ...
8:51ae333 feat: ship the /sys1grep Claude Code skill
Matching records are printed NUL-terminated too, so pipe them through tr '\0' '\n' to read them.
File names (-l) and counts (-c) stay on newlines, as they do in grep. With -z, -n numbers records,
-A / -B / -C count neighbouring records, and --chunk counts records per request.
One sentence at a time (--unit=sentence-by-*)
--unit=sentence-by-jev (or sentence-by-rule, below) judges each sentence instead of each line. The output is still lines, as in grep: every line a
matching sentence touches is printed, and on a terminal the sentence itself is in bold yellow (a regex match inside it, in grep's bold red).
Wrapped lines are joined before splitting, so a sentence that runs over several lines is judged as one.
tests/prose.txt wraps an English paragraph and a Japanese one:
$ ./sys1grep -n --unit=sentence-by-jev -e "the author admits they made a mistake" tests/prose.txt
1:I should have checked the input
2:before shipping, and that was my
3:mistake. Next time I will add a test
The sentence starts on line 1 and ends at mistake. on line 3; only that part is colored, not
Next time I will add a test, which is judged separately and does not match.
On a terminal, with both meanings and -C 3 for context, the colors show where each sentence starts and ends
inside a line: lines 3 and 9 are colored only up to the end of the matching sentence, and lines 4-7 are context (-):
-o prints only the matching sentences, one per line, as grep -o prints only the matching part.
-n then gives the line where the sentence starts. Japanese is joined without a space, as are Chinese,
Thai, Lao, Khmer, Myanmar and Tibetan, which do not put spaces between words:
$ ./sys1grep -n -o --unit=sentence-by-jev -e "the author admits they made a mistake" -e "customer is asking for a refund" tests/prose.txt
1:I should have checked the input before shipping, and that was my mistake.
8:先週買った掃除機が初日から動かないので返金してほしいです。
Without -o, -c counts lines and -A / -B / -C count lines, as usual. With -o they count sentences.
With -z, each record is split on its own and matching records are printed whole.
Jev finds a matching sentence inside a long line on its own, so --unit=sentence-by-jev is not needed for accuracy.
Use it to see which sentence matched, to get the sentences with -o, and when AND should hold within one
sentence: the expression is evaluated per sentence. For the same reason -v X alone prints every line
with at least one sentence that is not X; to find lines that are not X as a whole, leave --unit at line.
Japanese and Chinese entries often end without 。: a chat message, a support ticket, a memo line. Joining
them would glue separate entries into one "sentence". So with
--unit=sentence-by-jev sys1grep asks Jev about each unpunctuated break next to a script written without word
spaces: "does this line break end a sentence or entry, or is it a wrap inside a sentence?" It sends 30 lines
per request with one yes/no per break, and keeps the lines apart when the answer is 0.7 or more. On
tests/corpus.txt this keeps the four one-line Japanese tickets apart, so
--unit=sentence-by-jev finds the same refund requests (lines 14 and 18) as a line-by-line search, where the rules alone
merged the tickets and missed line 18. The extra requests cost about as much as one more meaning; use
--unit=sentence-by-rule to skip them. Breaks between English lines are never asked: joining them keeps a space,
and the full stop still ends the sentence.
Where a newline cannot be inside a sentence, lines are not joined: at a blank line, next to brackets or
; (JSON, code), and before a line starting with - * + # > " or a digit (list items,
headings, quotes, numbers, timestamps). So JSONL keeps one line per record and each line is split on its
own. Sentences are cut by Intl.Segmenter (Unicode UAX #29),
which splits at . ! ? 。 ! ? but also after abbreviations such as Mr.. Logs are not prose:
consecutive log lines that start with a letter, such as WARN ... after ERROR ..., get joined.
One function at a time (--unit=function)
In code the question is usually about a function ("retries on network failure"), and one line of it rarely
says so. --unit=function judges each function and prints its lines, with -n giving file line numbers:
$ ./sys1grep -n --unit=function -e "retries on network failure" src/net.js
src/net.js:3:async function fetchWithRetry(url) {
src/net.js:4: for (let i = 0; i < 5; i++) {
src/net.js:5: try { return await fetch(url); }
src/net.js:6: catch { await sleep(2 * i 100); }
src/net.js:7: }
src/net.js:8:}
A function runs from a funcname line to the line before the next one, as git grep -W finds it; lines
before the first funcname line are one unit, and blank lines at a function's end are not printed. The
funcname lines are git's: the diff= attribute in .gitattributes, then diff.
in git config. Without one, sys1grep has its own rule for JavaScript/TypeScript (a top-level function,
class, or const / let / var bound to a function) and Python (def / class, nested ones too); a decorator line (@retry) starts
the function it decorates. Otherwise it uses git's default: a line starting with a letter, _ or $. git's own builtin drivers
(diff=python and so on) are not read, so a driver without an xfuncname falls back the same way:
$ printf '*.go diff=golang\n' >> .gitattributes
$ git config diff.golang.xfuncname '^(func|type)[[:space:]]'
-M defaults to 8000 characters here, as with -z; a longer function is judged on its first 8000.
--chunk counts functions, -A/-B/-C and -c count lines, as with the sentence units. -z, -g and
-o are refused.
From one function to another (--step-to)
A question about code often comes in two steps: where does --summarize hand the lines to the tool, and then,
what happens when that fails. The "that" in the second step is what the first one found. Multi-step matching
answers both in one run: the expression before --step-to finds the start, sys1grep walks the calls from it,
and the expression after --step-to picks the functions it reaches that match. Jev judges the two
expressions; no score decides which call to follow.
$ sys1grep -e "--summarize hands the lines to the tool" --step-to "what happens when the tool fails to start or answer" sys1grep.mjs
A meaning may start with a dash, as the first one does, on any search and for -e, -a, -v and -Q: a value that starts with - is taken as the meaning unless it is an option (an option holds no space).
Each path prints as a tree: the hop (the number of calls from the start), file:line and the function's name,
: after an end and - after a function on the way, as grep marks a match and its context. A regex-only
start and end send nothing, which makes a free reachability search. In Python 3.6's json, the functions
that load reaches and that hold a raise:
$ cd /usr/lib64/python3.6/json && sys1grep -e '/^\sdef load\b/' --step-to '/\braise\b/' .py
sys1grep: walk: 1 at hop 0, 1 at hop 1, 3 at hop 2, 1 at hop 3, 1 at hop 4, 1 at hop 5; stopped: no new unit
0 __init__.py-274-load
1 __init__.py:302:loads
2 decoder.py:334:decode
3 decoder.py:345:raw_decode
4 scanner.py-65-scan_once
5 scanner.py:28:_scan_once
- The expressions. The expression before
--step-tois the start's, as in any search; every-e/-a/
-v / -Q after it is the end's. --step-to MEANING and --step-to '/RE/' are --step-to -e ...; the file
names follow. A step is a search, not a hop: one --step-to is two steps, whatever the hops between them.
- The edges. Without
--edgesthe unit is a function (--unit=function) and an edge is a call:name(in
def under a decorator). --edges=FILE walks any other relation instead,
one edge per line, FROM_FILE:LINEFROM_NAMETO_FILE:LINETO_NAME ; a line stands for the unit that
holds it. --reverse walks the edges backwards, callee to caller.
- The walk. Breadth first from every start at once. A function's hop is its shortest distance from any
--hops=N, --hops=M..N and --hops=M.. pick the
hops an end may be at (default 0..: a start that matches the end expression is an end at hop 0). A path that
reaches no end is left out. stderr says how many functions each hop reached and why the walk stopped: no new
function, or --hops.
- The cost. Jev judges every function against the start expression in one batch, then the functions reached
--hops against the end expression in another. --max-cost asks again before the second, counting the
first. --dry-run answers every question 0, so it walks from every function the start expression could hold
for and shows that bound.
- Not with
-z,-g,-o,-c,-l,-A/-B/-C,--rank,--summarizeor--dedup.-padds the
-q prints nothing and exits 0 when there is an end.
One line per template (--dedup)
Cost is proportional to the text sent, and machine-generated logs are mostly one skeleton with a different
id or number in it. --dedup masks ids, hashes, numbers, dates and times, paths and URLs, groups lines by
the result, sends one line per group and reuses its answer for the rest. What is sent is that line's
original text, and every line is printed as itself:
--dedup takes auto, always or never; a bare --dedup is --dedup=always. **Default is never, for
now** (#143): auto is held back until real-run stats say it should be the default. Every run still
estimates locally, before anything is sent, what folding every kind (the best case) would save, and prints
one stderr hint when that is worth at least twice the cost of --dedup's own question:
$ sys1grep -e "a request failed" app.log
sys1grep: 2040 units fold to at most 31 templates; --dedup=auto would save ~68 requests (~44k tokens)
--dedup=auto asks and folds only when that estimate says it pays; a file that barely repeats sends every
line, asking nothing extra. --verbose / --dry-run print the decision and its numbers for every value.
$ ./sys1grep --dedup -n -e "a request failed" app.log
1:worker request 3fa9c1e27b failed: connection reset
2:worker request 88d0e41a5c failed: connection reset
4:worker request 0b7f2a9e13 failed: connection reset
3 of 6 lines matched; 4 sent to Jev (2 folded by --dedup, ~235 input tokens / ~$0.000010 saved, 21%) in 2 requests, 907 input tokens, ~$0.000038
Whether a value may be folded depends on the meaning: a number decides "disk usage is above 90%", a time
decides "happened at night". So Jev is first asked, one small request per meaning, which kinds of value
could change a match, and those kinds are kept apart (one of the two requests above). With
-e "disk usage is above 90%" the same file keeps 95% and 12% apart and prints only 5:disk usage 95%.
Measured on real logs, with nothing kept: a 43,071-line system log folds into 550 templates (1.2% of the
bytes), install.log to 21.1%, a Claude Code transcript (jsonl) only to 56.8%. It is for machine-generated
logs; prose has no shared skeleton, and a meaning that reads a timestamp folds almost nothing. With -z or
--unit=sentence-by-* the records or sentences fold instead of lines.
After a search, the summary line says what the folding saved: (2 folded by --dedup, ~235 input tokens /
~$0.000010 saved, 21%) above. It estimates what the folded lines would have cost as requests of their own and subtracts
what the value questions cost. On a small file or on prose, the result can be negative.
To see how far your own logs fold before paying for a search, node scripts/dedup-measure.mjs FILE...
counts lines, templates and the share of bytes sent, offline, with the same masks; --keep=num,time shows
a meaning that keeps those kinds apart.
The request's other lines are each line's context (#9), and --dedup changes them, so a line near the
threshold can be judged differently than in a full pass.
A summary instead of the lines (--summarize)
When many lines match, what you want is often the gist: what they say about the meaning you searched for.
--summarize pipes what sys1grep would print to claude -p --model haiku (no tools, no settings, no CLAUDE.md),
asks it to summarize the lines as they bear on the meanings, and prints its answer instead of the lines.
$ sys1grep -r -n --summarize -e "the API key is read from a file" .
The key is read in sys1grep.mjs:22 from ~/.config/sys1grep/.env, never from ./.env (sys1grep.mjs:18), ...
The expensive model reads only what Jev kept. Asking why ./.env is no longer read of this repository's
git log (168 commits), Claude's input fell from 22,059 tokens to 1,356, the total cost with Jev's from
$0.094 to $0.013, with the same answer (#69). It pays when the answer sits in a few lines.
- The matching lines leave the machine a second time, to Anthropic (or whatever TOOL talks to).
-n,-A/-B/-C,-pand file names go in as they would print; colors never do. No match runs nothing (exit 1).SYS1GREP_SUMMARIZERpicks the TOOL of a bare--summarize,SYS1GREP_SUMMARIZER_MODELits model.-q,-land-cprint no lines, so they cannot be combined with it.- In
SYS1GREP_OPTSit summarizes every search, so every match goes to TOOL's provider too;
--no-summarize turns it off for one search.
- The answer is plain text: an LLM that is not told writes Markdown, noise on a terminal, so plain is asked for.
--format=markdown or =html asks for those instead (… --format=html … > summary.html).
The answer prints as it comes, unchecked. In SYS1GREP_OPTS it is a standing preference, ignored without --summarize or --rank.
- With
--dedup, the TOOL gets what Jev got: each template's representative once, marked(×N like it),
- More than 200 KB (about 50k tokens) is not sent at all: exit 2, with the size, before the TOOL is paid.
--summarize-prompt=TEXT adds your own instruction after the fixed one (how long, what to focus on):
$ sys1grep -r -n --summarize --summarize-prompt="3 lines or fewer, just which file to fix" \
-e "the API key is read from a file" .
It needs --summarize; empty TEXT is the same as leaving it out. It comes after the --format sentence,
so TEXT can override the format.
Other TOOLs:
$ sys1grep --summarize=llm -e "..." FILE # Simon Willison's llm, tools off (no -T given)
$ sys1grep --summarize=pi -e "..." FILE # pi --print --no-tools --no-session ...
$ SYS1GREP_SUMMARIZER_MODEL=qwen3.5:9b sys1grep --summarize=ollama -e "..." FILE
$ SYS1GREP_SUMMARIZER_MODEL=some-id sys1grep --summarize=lmstudio -e "..." FILE # model id from GET /v1/models
$ SYS1GREP_SUMMARIZER_MODEL=some-id sys1grep --summarize=http://localhost:8080/v1 -e "..." FILE # llama.cpp, vLLM, LocalAI, a gateway
ollama and lmstudio talk to a local OpenAI-compatible server (POST /v1/chat/completions) by fetch, no
CLI: the matching lines never leave the machine a second time. A plain http(s):// URL is any other
OpenAI-compatible server. All three need SYS1GREP_SUMMARIZER_MODEL: none has a default model.
SYS1GREP_SUMMARIZER_API_KEY goes as Authorization: Bearer to a URL TOOL only (never SYS1GREP_API_KEY,
which is Jev's). OLLAMA_HOST moves ollama's host, as it does for the ollama CLI itself.
Best first (--rank)
Matches print in file order, as grep prints them. With many, --rank prints the results best first, each under a
numbered header. A result is a match with its -A/-B/-C lines; matches whose context touches are one result.
$ sys1grep -n -C1 --rank -e "the refund was refused" tickets/
- tickets/b.txt
tickets/b.txt-2-Order #1234, placed 2026-08-01.
tickets/b.txt:3:Your refund was declined: the order is older than 30 days.
tickets/b.txt-4-Please contact support.
- tickets/a.txt
tickets/a.txt-11-Order #88 arrived.
tickets/a.txt:12:Refund request received.
tickets/a.txt-13-We will check it.
--rankis--rank=jev: after the search, Jev is asked of each result, lines and context together, whether it
--chunk results to a request;
--dry-run / -i show an upper bound, since which results there are is known only after the search.
--rank=matchsorts by each result's highest match probability, with no request. It is the answer to a yes/no
-pputs the score on the header (1. [0.96] tickets/b.txt).-llists the files by their best result.--format=markdownwrites a## 1. tickets/b.txtheading and a fenced block per result;--format=htmlone
--summarizegets the results in ranked order; with--dedupa representative is one result.- It needs a