Profile
Back to NewsBack
GitHub Trending 40 min
Reader Mode
uehaj/sys1grep: grep by meaning, across languages. TypeSafe Jev scores every line against a meaning; combine meanings with AND/OR/NOT. 意味で探す grep。日本語で英語を、英語で日本語を検索できる

uehaj/sys1grep: grep by meaning, across languages. TypeSafe Jev scores every line against a meaning; combine meanings with AND/OR/NOT. 意味で探す grep。日本語で英語を、英語で日本語を検索できる

6 hours ago

sys1grep — System 1 for your files; uses TypeSafe AI's Jev

日本語版はこちら

Background and design notes (Japanese): Jevのキラーアプリ、「意味で探す grep」を作った on Zenn.

Renamed from semgrep (0.5.0). Two reasons. The name collided with the static-analysis tool Semgrep, a trademark. And the tool was never meant for Jev alone: it targets System 1 models in general, and I'd like it to work with any models TypeSafe AI may release after Jev under other names :). So from 0.5.0 this is sys1grep (System 1 + grep): npm package @uehaj/sys1grep (was @uehaj/semgrep), commands sys1grep / git sys1grep, environment SYS1GREP_, config ~/.config/sys1grep/.env. The old SEMGREP_ names and ~/.config/semgrep/.env still work for one minor release, each use printing a deprecation line; 1.0.0 drops them. Details: CHANGELOG. The Zenn article's URL keeps its old slug.

A grep that finds lines by what they mean, not by regular expressions. Matching is done by Jev, the System One model from TypeSafe AI. Jev does not generate text. It answers typed questions with probabilities, so for every line sys1grep asks "does this line match the meaning network failure?", gets a probability back, and applies a threshold.

./sys1grep -n -e "customer is angry or frustrated" tickets.txt

sys1grep demo: a Japanese meaning finds refund requests in six languages; " style="max-width:100%;">

▶ Click the demo, or open uehaj.github.io/sys1grep, for the full demo on the landing page.

  • Zero dependencies. One file, Node.js 20.16+ and fetch.
  • Fast. 30 lines go into one request, requests run 8 at a time. A 210-line file finishes in under a second.
  • Meanings combine with AND / OR / NOT.
  • Language-agnostic. The meaning and the text can each be in any language. A Japanese meaning finds French, Russian, Chinese and Korean lines alike. No translation step, same speed, same cost.

Search across languages

The meaning and the text do not have to share a language. Jev compares concepts, not words, so one query finds matching lines in every language the file contains.

An English meaning finds Japanese lines. None of the hits contain "angry" or "frustrated", and two of them are in Japanese:

$ ./sys1grep -n -e "customer is angry or frustrated" tests/corpus.txt
14:ユーザー山田さんからの問い合わせ: 返金してほしい、商品が壊れていた
16:ユーザー佐藤さんからの問い合わせ: 注文した覚えのない請求が来ています。至急確認してください
18:I want my money back. The item arrived broken and customer service ignored me.
21:Your product ruined my weekend. Never buying from you again.
23:This is the third time I'm writing. Nobody has replied to my previous emails.

A Japanese meaning finds English lines, with the same confidence as the Japanese ones:

$ ./sys1grep -n -p -e "返金の要求" tests/corpus.txt
14:ユーザー山田さんからの問い合わせ: 返金してほしい、商品が壊れていた	[0.97]
18:I want my money back. The item arrived broken and customer service ignored me.	[0.95]

It is not a Japanese/English feature. tests/multi.txt holds refund requests and thank-you notes in French, Russian, German, Spanish, Chinese and Korean. One Japanese meaning finds all six refund requests; a Russian meaning does the same:

$ ./sys1grep -n -p -e "顧客が返金を求めている" tests/multi.txt
1:Je veux être remboursé, le produit est arrivé cassé.	[0.98]
3:Я требую вернуть деньги, товар не работает.	[0.97]
5:Ich möchte mein Geld zurück, das Gerät ist defekt.	[0.97]
7:Quiero un reembolso, el paquete llegó vacío.	[0.97]
9:我要求退款,商品坏了。	[0.97]
11:환불해 주세요. 제품이 고장났어요.	[0.97]

$ ./sys1grep -n -p -e "клиент требует возврат денег" tests/multi.txt 1:Je veux être remboursé, le produit est arrivé cassé. [0.94] 3:Я требую вернуть деньги, товар не работает. [0.97] 5:Ich möchte mein Geld zurück, das Gerät ist defekt. [0.92] 7:Quiero un reembolso, el paquete llegó vacío. [0.89] 9:我要求退款,商品坏了。 [0.92] 11:환불해 주세요. 제품이 고장났어요. [0.89]

This makes sys1grep useful for mixed-language logs and ticket dumps, and for teams whose members query in different languages. One caveat: TypeSafe documents English as the most accurate language, and in our tests Japanese meanings wobble a little more near the threshold. When a query is borderline, phrasing the meaning in English is the safer choice.

How this differs from vector search

If all you want is "lines about X", embedding similarity gives you the same lines. sys1grep differs in what it judges: not how close a line is to a topic, but whether a proposition holds for that line. Jev reads the line and the question together (a cross-encoder shape), so who did what, negation, and "asked for" versus "already done" all change the answer. An embedding of the line is fixed before it ever sees your query, so it can only measure topical closeness.

All eight lines below are "about a refund" (the same contrast as tests/contrast.txt, translated to English as tests/contrast.en.txt — see Search across languages above for the case where the meaning and the text differ in language). Only two are a customer asking for one:

$ ./sys1grep -n -p -t 0 -e "customer is asking for a refund" tests/contrast.en.txt
1:I want a refund. The item was broken.	[0.99]
2:The refund has been processed. Please check your account.	[0.11]
3:Our refund policy is within 30 days of purchase.	[0.09]
4:The manager denied the refund request yesterday	[0.17]
5:I demand a full refund immediately	[0.93]
6:Refunds are processed within 5 business days	[0.07]
7:The support agent got angry and hung up the phone.	[0.05]
8:The customer got angry and hung up the phone.	[0.06]

Because each meaning yields an independent probability, **logical AND and NOT are plain boolean operations**, not a trick with set differences or "negative queries":

# about a refund, but NOT a customer asking for one → completed, policy, denied, timelines
$ ./sys1grep -n -e "about a refund" -v "the customer is asking for a refund" tests/contrast.en.txt
2:The refund has been processed. Please check your account.
3:Our refund policy is within 30 days of purchase.
4:The manager denied the refund request yesterday
6:Refunds are processed within 5 business days

angry AND it is the customer, not the staff

$ ./sys1grep -n -e "someone is angry" -a "the customer, not the staff, is the one acting" tests/contrast.en.txt 5:I demand a full refund immediately 8:The customer got angry and hung up the phone.

Line 7, The support agent got angry and hung up the phone., scores 0.05 on the second meaning and is excluded, even though it differs from line 8 by one word (support agent vs. customer) — wording an embedding would place right next to line 8's. (We could not reproduce a cosine-similarity number for this pair: no embedding-model access was available in this environment, and no prior measurement script exists in the repository or its history to rerun. The argument stands on the wording alone: an embedding of line 7 has no way to see that "support agent" changes who the sentence is about.)

Two more practical consequences. The probabilities are calibrated, so one threshold (0.5) works across queries, where cosine scores need top-k or per-query tuning. And there is no index to build: sys1grep reads the files in front of you. The flip side is that every query pays for the whole corpus again, so for repeated queries over a large, fixed corpus a vector index is cheaper and faster.

Regex terms

-e/-a/-v '/pattern/flags' (first and last character /, JavaScript flags) is matched locally, as a plain regex, with no request at all. It prefilters its AND term: only the lines it holds for ever ask that term's meanings, so a cheap regex in front of a meaning cuts both the bill and the wait. A line is sent only if some term's regexes all hold for it (a term with no regex holds for every line), so with -e '/re/' -a A -e B a line without re is still sent, asked B only. A query of regex terms alone sends nothing, even with --unit=sentence-by-jev: the rules alone then decide where wrapped lines join. Anything that isn't shaped like /…/flags is still a meaning, so -e '/etc 以下のファイルを変更している' (no closing /) is unaffected; a meaning that really starts and ends with / can be written with a leading space to dodge the regex reading.

$ ./sys1grep -e '/ERROR|FATAL/' app.log                               # no requests at all
$ ./sys1grep -e '/timeout/i' -a '顧客に影響が出ている' app.log         # only lines with "timeout" go to Jev

A regex's named and numbered groups pass to the other meanings of the same AND term as $, $1-$99, $&, $$ — ECMAScript's replacement-pattern syntax (String.prototype.replace's GetSubstitution), with one deviation: $ naming no group is an error rather than an empty string (so is a reference to a negated regex's group). A $n naming no group stays literal, as in ECMAScript, so $100 以上の請求 is unaffected.

$ ./sys1grep -e '/(?<date>\d{4}-\d\d-\d\d) (?<time>\d\d:\d\d)/' \
            -a '$<time> が深夜(0時〜5時)であり、$<date> が週末である' app.log

2026-09-19 03:12 ... → asks "03:12 が深夜(0時〜5時)であり、2026-09-19 が週末である"

Prefer $ and single quotes: $ survives double quotes in sh/bash/zsh; $1, $time and ${time} don't (the shell expands them itself). -p prints 1.00/0.00 for a regex term.

On a terminal every match of a regex is in grep's match color (bold red), for the terms that held; a negated regex is never colored. -o prints each match on a line of its own, as grep -o; a line only meanings matched prints whole, since a meaning has no matching part.

$ ./sys1grep -o -n -e '/[A-Z]+-\d+/' -a 'the ticket is still open' notes.txt   # the ticket ids, one per line

Sending less

Every line sent to Jev costs money and time, so the cheapest line is the one never sent. From the widest cut to the narrowest:

  • Which files. -r skips .git, node_modules, binary files, likely secrets, generated files (source maps,
minified JS/CSS, lock files) and what git ignores; git sys1grep searches tracked files only. --include / --exclude (file-name globs) and --changed-within (30m, 7d, today, this-week, a date) narrow them further. A meaning that restricts itself to a language or a time of change narrows them by itself: see Scope from the meaning.
  • Which lines. A regex term is matched locally, and only the lines it holds for are asked
its AND term's meanings. Blank lines are never sent. Only the first -M/--max-columns characters of a line (default 2000, 8000 with -z) are sent: it is still searched and judged, just not past that cutoff.
  • How many times. --dedup judges one line per template: lines that
differ only in ids, numbers, times or paths share one answer. auto/always/never, off by default for now: a hint says when --dedup=auto would pay.
  • Check before paying. --dry-run sends nothing and prints the settings the search would run with, the
files, how many lines each would send and every request with its questions. Its last line estimates the input tokens and, for TypeSafe itself, the price (~3178 input tokens, ~$0.000133; within about 10%). -i shows the same totals on the terminal and sends only after y. Before the bulk of requests, every target is sized (--max-filesize, default 10M); one over it is skipped outright, like rg's own --max-filesize, named on stderr (-y does not affect it). A .gz is sized decompressed (zlib stops at the limit, so an oversized one is never inflated in full). stdin is sized once it is read (it is already in memory), and skipped the same way if it is over; a --cached or : target is a blob rather than a file, sized from its content (already read in one git cat-file --batch call, not a fresh git process per blob); -g's commits stay out of this, since each is already bounded by -M when it is sent. The input about to be sent is priced too (--max-cost, default 1 USD); over it, one question asks to continue on the terminal, -y answers it yes, and without a terminal it is exit 2 (this applies with -q too — scripts pass -y). With -r or git sys1grep, a term no regex narrows that is about to send more than 10,000 lines gets one stderr line with the same totals, naming the term (sys1grep: sending 214,913 of 231,502 lines from 1,247 files (~9.1M input tokens, ~$0.38); the term "…" has no regex to narrow it. Add -a '/RE/' to it, …). The search goes on; -q silences the line, -y does not.
$ sys1grep --dry-run -r --include='*.log' --changed-within=today -e '/ERROR|FATAL/' -a 'a customer is affected' logs/

--verbose (or --dry-run) also shows where a setting that was not typed on the command line came from (SYS1GREP_OPTS, an environment variable, ~/.config/sys1grep/.env, or a preset's default), so a result that surprises you can be traced back to its source. The key's value never appears, only which option or variable supplied it:

$ SYS1GREP_OPTS='--level strict' sys1grep --verbose -e "the API key is read from a file" .
sys1grep: endpoint api.typesafe.ai/v1/systemone (default), model jev-latest (default)
sys1grep: key: SYS1GREP_API_KEY (~/.config/sys1grep/.env)
sys1grep: SYS1GREP_OPTS: --level strict
sys1grep: options: --level strict (SYS1GREP_OPTS) = -t 0.7 -T 0.3, --chunk 30, -j 8, scope on, -M 2000, --max-filesize 10M, --max-cost 1
sys1grep: file ./a.py: 120 lines, 120 to send
…

Install

Two ways to use it: as a command-line tool (this section), or as a Claude Code skill (see Use it from Claude Code below). The skill falls back to npx @uehaj/sys1grep, so if you only use it through Claude Code you can skip the install here entirely and just set the API key.

Requires Node.js 20.16 or later. No other dependencies.

npm install -g @uehaj/sys1grep
sys1grep --help

To try it without installing, run it through npx (the first run downloads the package, later runs use the cache):

npx @uehaj/sys1grep -n -e "customer is angry or frustrated" tickets.txt

Then give it an API key from the TypeSafe console. Any one of these works:

export SYS1GREP_API_KEY=your-key                       # environment variable
echo 'SYS1GREP_API_KEY=your-key' > ~/.config/sys1grep/.env    # per user (mkdir -p first)

Variables already in the environment win; otherwise ~/.config/sys1grep/.env fills them in. A .env in the current directory is never read: it may belong to a repository you just cloned, and could send your key elsewhere through SYS1GREP_URL. For per-project settings, load a file yourself: node --env-file=.env "$(command -v sys1grep)" .... TYPESAFE_API_KEY is accepted too when SYS1GREP_API_KEY is not set.

Default options

SYS1GREP_OPTS holds options to apply on every call, read like the settings above. It is split on spaces and put in front of the command line, so the command line wins: a later value counts, and --no-X turns a default flag off.

export SYS1GREP_OPTS='--level strict -j 8 -n'
sys1grep -e "payment failed" app.log                  # strict, 8 at once, line numbers
sys1grep --level loose --no-n -e "payment failed" app.log

A script calling sys1grep would pick these up too (grep dropped GREP_OPTIONS for that reason). Call it as SYS1GREP_OPTS= sys1grep ... in scripts.

Other endpoints

The API is configured by exactly three settings: SYS1GREP_API_KEY (or TYPESAFE_API_KEY), SYS1GREP_URL and SYS1GREP_MODEL. Any endpoint that speaks TypeSafe's POST /v1/systemone works. The key is sent to SYS1GREP_URL as is, so set the two together. On the command line, --sys1-model=ID, --sys1-url=URL and --sys1-api-key=KEY override the three. A key given this way shows up in ps and shell history, so prefer ~/.config/sys1grep/.env for it.

# OpenRouter
SYS1GREP_URL=https://openrouter.ai/api/v1/systemone SYS1GREP_API_KEY=sk-or-... sys1grep -e ...

Vercel AI Gateway

SYS1GREP_URL=https://ai-gateway.vercel.sh/typesafe/v1/systemone SYS1GREP_MODEL=typesafe-ai/jev SYS1GREP_API_KEY=vck_... sys1grep -e ...

A compatible server that needs no key: no Authorization header is sent

SYS1GREP_URL=http://localhost:8000/v1/systemone sys1grep -e ...

The summary line shows the cost the endpoint reports (usage.cost), or for TypeSafe itself an estimate at list price marked ~.

From source: git clone https://github.com/uehaj/sys1grep.git && cd sys1grep && npm install -g ., or run it in place with node sys1grep.mjs ....

Examples

All examples run against tests/corpus.txt, a 51-line mix of server logs, support tickets in English and Japanese, source code, SQL and small talk. Where a Japanese line would otherwise show up in the output below, this section instead uses tests/corpus.en.txt, the same 51 lines with the Japanese ones translated to English — the cross-language behaviour itself is shown once, in Search across languages above.

Find lines by a concept, in any language

$ ./sys1grep -n -e "customer is angry or frustrated" tests/corpus.en.txt
14:Inquiry from user Yamada: I want a refund, the item was broken
16:Inquiry from user Sato: I'm being charged for an order I never placed, it looks like fraud, please check urgently
18:I want my money back. The item arrived broken and customer service ignored me.
21:Your product ruined my weekend. Never buying from you again.
23:This is the third time I'm writing. Nobody has replied to my previous emails.
5 of 51 lines matched; 51 sent to Jev in 2 requests, 3054 input tokens, ~$0.000128

None of these lines contain the words "angry" or "frustrated".

Lines that answer a question (-Q)

-e asks whether a line states a meaning. -Q QUESTION (--question) finds lines that answer it instead, so you can write the question as you would ask it:

$ echo "The job failed." | ./sys1grep -e "Did the job succeed?"
$ echo "The job failed." | ./sys1grep -Q "Did the job succeed?"
The job failed.

As a meaning, "Did the job succeed?" is not what the line says (0.08), so -e finds nothing. As a question, the line answers it: no, it failed (0.83). A line that asks the question is close to it in meaning but answers nothing, and for a yes/no question a line that denies it still answers it:

$ ./sys1grep -n -Q "whether the server is down" tests/intent.txt
5:The server is down.
6:The server is healthy and responding normally.
2 of 17 lines matched; 17 sent to Jev in 1 request, 1087 input tokens, ~$0.000046

Both the confirming line and the denying line match: each settles whether the server is down. Is the server down? does not match, because asking is not answering; -e "asking whether the server is down" would match it instead. -Q X is shorthand for -e "the line answers: X", so it combines with -a / -v / ! and OR's with other terms exactly like -e.

OR: two meanings, and see the probabilities with -p

$ ./sys1grep -n -p -e "customer is asking for a refund" -e "delivery address change request" tests/corpus.en.txt
14:Inquiry from user Yamada: I want a refund, the item was broken	[0.99 0.01]
17:Inquiry from user Takahashi: I'd like to change the delivery address	[0.01 0.99]
18:I want my money back. The item arrived broken and customer service ignored me.	[0.96 0.01]
22:Can I change the delivery address for order #8821?	[0.02 0.99]
4 of 51 lines matched; 51 sent to Jev in 2 requests, 4275 input tokens, ~$0.000180

The bracket shows one probability per meaning, in the order given. Use it to pick a threshold.

With --color (on by default in a terminal) the probabilities are colored against the thresholds: green at or above -t, red below -T, yellow in between. Line numbers and file names use grep's colors.

!colored output: line numbers in green, probabilities in green or red

AND NOT: network errors, excluding retries

$ ./sys1grep -n -e "network or remote connection failure" -v "a retry is happening or was attempted" tests/corpus.txt
4:2026-09-19 08:02:30 ERROR connection reset by peer while calling payment-gateway
6:2026-09-19 08:02:35 ERROR timeout after 5000ms waiting for payment-gateway
9:2026-09-19 08:10:44 ERROR DNS lookup failed for api.example.com
11:2026-09-19 09:00:00 ERROR SSL handshake failed: certificate expired
13:network unreachable: no route to host 10.0.0.5
30:except ConnectionError as e:
31:    logger.error("upstream unreachable: %s", e)
7 of 51 lines matched; 51 sent to Jev in 2 requests, 5112 input tokens

Line 5, retrying payment-gateway request (attempt 2/3), is a network failure but is dropped by -v.

Mixed: (finance AND negative) OR weather

$ ./sys1grep -n -e "about economy, finance or markets" -a "the news is negative or a decline" -e "about weather" tests/corpus.en.txt
36:Today's weather is sunny, high of 28 degrees
43:Stock prices fell 3% after the earnings report missed expectations.
48:Tomorrow's forecast is rain so I'll bring an umbrella
3 of 51 lines matched; 51 sent to Jev in 2 requests, 5499 input tokens, ~$0.000231

The central bank raised interest rates is about finance but not a decline, so it is out.

Strictness presets

$ ./sys1grep --level strict -n -e "a security risk or dangerous destructive operation" tests/corpus.en.txt
33:DROP TABLE sessions;
49:API keys must never be committed to the repository.

$ ./sys1grep --level loose -n -e "a security risk or dangerous destructive operation" tests/corpus.en.txt 11:2026-09-19 09:00:00 ERROR SSL handshake failed: certificate expired 16:Inquiry from user Sato: I'm being charged for an order I never placed, it looks like fraud, please check urgently 33:DROP TABLE sessions; 49:API keys must never be committed to the repository.

strict keeps only what the model is sure about. loose also pulls in the expired certificate and the suspicious-billing ticket.

Recursive search and file names only

$ ./sys1grep -r -n -e "customer is asking for a refund" tests/tickets/
tests/tickets/a.txt:7:Ticket #16: I want a refund, the item was broken.
tests/tickets/sub/b.txt:1:The customer wants a refund for the broken lamp.

$ ./sys1grep -rl -e "customer is asking for a refund" tests/tickets/ tests/tickets/a.txt tests/tickets/sub/b.txt

-r walks directories in sorted order and skips .git, node_modules, .ssh, .aws, .gnupg, .kube, .docker, binary files (a NUL byte in the first 8 KB, or a PDF; UTF-16 with a BOM is read as text), files that usually hold secrets (.env, .netrc, .npmrc, .pypirc, .pgpass, .git-credentials, .pem, .key, .p12, .pfx, .jks, .keystore, id_rsa and friends; names compared without case) and generated files (.map, .min.js, *.min.css, package-lock.json, yarn.lock, pnpm-lock.yaml, Cargo.lock, poetry.lock, composer.lock, Gemfile.lock, go.sum). **Every line that is searched is sent to the TypeSafe API**, so point -r at a directory you mean to scan. Inside a git repository, -r also skips what git ignores (.gitignore, .git/info/exclude, the global excludes file), so build output and local files stay home; a tracked file is searched even if it matches. A file named *.gz (a rotated log: app.log.1.gz) is read decompressed, as zgrep does, and printed by its name on disk (app.log.1.gz:12:...); --include='*.gz' picks those, the binary sniff sees the decompressed bytes, a gzipped secret (private.key.gz) is skipped like the plain one, and a corrupt .gz is an unreadable file. Only gzip, only by the name. A file named explicitly on the command line is always searched, even if it matches the skip list or is ignored; so is a directory that git ignores, when you name it (sys1grep -r -e ... dist). -l prints each matching file once, in the order matches are found, and works with or without -r. -c prints the number of matching lines per file instead.

Scope from the meaning

A meaning that restricts its matches to some kind of file can only match in such files. With -r and git sys1grep, each meaning first asks Jev, in one small request, a yes / no per candidate, and a yes at 0.6 or more narrows the files before anything else is sent. The rest is never read or sent. Each scope is reported on stderr with Jev's answer, so a wrong one is visible:

$ sys1grep -r -e 'Python でリトライ処理を書いている箇所' .
sys1grep: scope: .py .pyi *.pyw (from "Python files: 0.94")
sys1grep: scope: 12 of 340 files
$ sys1grep -r -e '昨日変えた箇所で認証を扱っている' src/
sys1grep: scope: modified since 2026-09-25 00:00 (from "what was changed yesterday: 0.91")
sys1grep: scope: 3 of 120 files
  • Language or format: 26 candidates (Python, JavaScript, TypeScript, Go, Rust, Java, Kotlin, Ruby, PHP, C, C++,
C#, Swift, Scala, R, shell script, SQL, HTML, CSS, Markdown, YAML, JSON, TOML, XML, Dockerfile, Makefile), with GitHub Linguist's extensions and file names. Several yes answers are alternatives ("JavaScript か TypeScript").
  • Time of change: 14 spans: the last minute, the last hour, today, yesterday, the last 1 / 2 / 3 / 7 / 30 days,
this month, last week, last month, the last year, this fiscal year (from April 1). The narrowest span answered yes is taken, by its start only: a file changed yesterday may have been modified again today.
  • Place: test code (tests/ test/ __tests__/ spec/ e2e/, test_.py _test. .test. .spec. Test.java
_spec.rb, and every .rs: Rust keeps unit tests inside the file), database migrations (migrat, Flyway's V1__.sql), the README, the changelog (CHANGELOG CHANGES HISTORY NEWS), documentation (.md .rst .adoc *.txt, docs/), source code (whatever is not documentation, so languages missing from the dictionary still count) and logs (.log .log.N .out .err, logs/), by the path conventions of JS, Python, Go, Java, Ruby, Rust and PHP. Several places are alternatives ("README か CHANGELOG に").
  • git: in a repository, a time goes by commits: a committed file needs a commit at or after the start (the
committer date; a checkout sets every mtime to now), an uncommitted one its mtime. States: uncommitted (worktree and index against HEAD, and untracked files), staged, untracked, this branch (since it left origin/HEAD, main or master), not pushed (@{upstream}..HEAD, else commits on no remote branch), and mine (user.email's commits and the uncommitted files). Authors: the 30 with the most commits (git shortlog), each a candidate by name and e-mail; a yes narrows to every file a commit of theirs touched. Outside a repository none of these is asked, and a time goes by the mtime.
  • Jev reads the whole meaning, in any language: "案A、B、Cで比較" is not about C files, and a date quoted in a
comment is not when the file changed. Nothing is extracted from the text; the candidates are fixed.
  • Per term. -e A -e B still searches B in the files A's scope leaves out; within an AND term, and across
categories ("Python のテストコード"), the scopes intersect. Negated meanings (-v, !) are not asked. The meaning is sent unchanged.
  • The question costs one small request per meaning, sent only when -r / git sys1grep found something to
narrow, and after -i's answer. --dry-run shows it as [scope]. As with --include, with -r a file named on the command line and stdin are never narrowed, and git sys1grep's pathspecs are narrowed like the rest. --no-auto-scope turns it off (--auto-scope turns it back on).

As a git subcommand (git sys1grep)

npm install -g also installs git-sys1grep, so git runs it as git sys1grep. Like git grep, it searches only the files git tracks (ignored files and build output are never sent), and FILE arguments are pathspecs relative to the current directory.

$ cd tests && git sys1grep -l -e "customer is asking for a refund" fixture.txt tickets
fixture.txt
tickets/a.txt
tickets/sub/b.txt

Without FILE it searches every tracked file under the current directory. The -r skip list (.env*, keys, ...) applies even to tracked files. --include, --exclude and --changed-within filter the tracked files, those named by a pathspec included. --changed-within reads the working tree's modification times, not git history: right after a clone or a checkout, every file it wrote counts as just changed. For help use git sys1grep -h: git takes --help itself and looks for a man page.

--cached, --untracked and a ... choose what is searched instead of the working tree, as git grep has them (only one of the three at a time):

$ git sys1grep --cached -e "a retry is attempted"       # staged, uncommitted changes included
$ git sys1grep --untracked -e "a retry is attempted"    # tracked files plus untracked ones (.gitignore still applies)
$ git sys1grep -n -e "a retry is attempted" main v0.3.1 -- '*.py'
main:src/job.py:42:    retry(job, times=3)
v0.3.1:src/job.py:40:    retry(job)

A (a branch, tag, commit or @{u}) is any argument before -- that resolves as a revision; several may be given, searched in the order given, each line prefixed with the name as typed (@{u}:path, not the branch it resolves to). --changed-within needs the working tree: a blob (--cached, a ) has no mtime of its own. A 's own pathspec takes a glob too, same as elsewhere in sys1grep (git diff-tree against the empty tree, not git ls-tree's own literal/directory-prefix match); --include / --exclude (by name) still work. An old can hold a secret an ordinary file once carried and was later removed from: the skip list drops files by name only, not by what changed since.

Everything that is not something

./sys1grep -v "a timestamped server log line" mixed.txt   # like grep -v
./sys1grep -e "source code or SQL" -v "SQL" src.txt       # code, but not SQL
cat app.log | ./sys1grep -e "the deploy failed or was rolled back"

Records that span several lines (-z)

The unit of judgement is a line. That is right for logs and source, and wrong when one record spans several lines. -z makes the unit a NUL-terminated record instead, exactly as in grep -z, so it pairs with the tools that already emit records: git log -z, find -print0, xargs -0.

A proposition like "this commit changes user-visible behaviour" is true of a whole commit, not of any one line in it:

$ git log -z --format='%h %s %b' | ./sys1grep -z -n -e "the change alters user-visible behaviour" -v "documentation only"
7:21120e9 Revert "feat: ship the /sys1grep Claude Code skill" ...
8:51ae333 feat: ship the /sys1grep Claude Code skill

Matching records are printed NUL-terminated too, so pipe them through tr '\0' '\n' to read them. File names (-l) and counts (-c) stay on newlines, as they do in grep. With -z, -n numbers records, -A / -B / -C count neighbouring records, and --chunk counts records per request.

One sentence at a time (--unit=sentence-by-*)

--unit=sentence-by-jev (or sentence-by-rule, below) judges each sentence instead of each line. The output is still lines, as in grep: every line a matching sentence touches is printed, and on a terminal the sentence itself is in bold yellow (a regex match inside it, in grep's bold red). Wrapped lines are joined before splitting, so a sentence that runs over several lines is judged as one. tests/prose.txt wraps an English paragraph and a Japanese one:

$ ./sys1grep -n --unit=sentence-by-jev -e "the author admits they made a mistake" tests/prose.txt
1:I should have checked the input
2:before shipping, and that was my
3:mistake. Next time I will add a test

The sentence starts on line 1 and ends at mistake. on line 3; only that part is colored, not Next time I will add a test, which is judged separately and does not match.

On a terminal, with both meanings and -C 3 for context, the colors show where each sentence starts and ends inside a line: lines 3 and 9 are colored only up to the end of the matching sentence, and lines 4-7 are context (-):

!--unit=sentence-by-jev -C 3 --color: the matching sentences in bold yellow, up to mistake. on line 3 and 返金してほしいです。 on line 9; lines 4 to 7 as context

-o prints only the matching sentences, one per line, as grep -o prints only the matching part. -n then gives the line where the sentence starts. Japanese is joined without a space, as are Chinese, Thai, Lao, Khmer, Myanmar and Tibetan, which do not put spaces between words:

$ ./sys1grep -n -o --unit=sentence-by-jev -e "the author admits they made a mistake" -e "customer is asking for a refund" tests/prose.txt
1:I should have checked the input before shipping, and that was my mistake.
8:先週買った掃除機が初日から動かないので返金してほしいです。

Without -o, -c counts lines and -A / -B / -C count lines, as usual. With -o they count sentences. With -z, each record is split on its own and matching records are printed whole.

Jev finds a matching sentence inside a long line on its own, so --unit=sentence-by-jev is not needed for accuracy. Use it to see which sentence matched, to get the sentences with -o, and when AND should hold within one sentence: the expression is evaluated per sentence. For the same reason -v X alone prints every line with at least one sentence that is not X; to find lines that are not X as a whole, leave --unit at line.

Japanese and Chinese entries often end without 。: a chat message, a support ticket, a memo line. Joining them would glue separate entries into one "sentence". So with --unit=sentence-by-jev sys1grep asks Jev about each unpunctuated break next to a script written without word spaces: "does this line break end a sentence or entry, or is it a wrap inside a sentence?" It sends 30 lines per request with one yes/no per break, and keeps the lines apart when the answer is 0.7 or more. On tests/corpus.txt this keeps the four one-line Japanese tickets apart, so --unit=sentence-by-jev finds the same refund requests (lines 14 and 18) as a line-by-line search, where the rules alone merged the tickets and missed line 18. The extra requests cost about as much as one more meaning; use --unit=sentence-by-rule to skip them. Breaks between English lines are never asked: joining them keeps a space, and the full stop still ends the sentence.

Where a newline cannot be inside a sentence, lines are not joined: at a blank line, next to brackets or ; (JSON, code), and before a line starting with - * + # > " or a digit (list items, headings, quotes, numbers, timestamps). So JSONL keeps one line per record and each line is split on its own. Sentences are cut by Intl.Segmenter (Unicode UAX #29), which splits at . ! ? 。 ! ? but also after abbreviations such as Mr.. Logs are not prose: consecutive log lines that start with a letter, such as WARN ... after ERROR ..., get joined.

One function at a time (--unit=function)

In code the question is usually about a function ("retries on network failure"), and one line of it rarely says so. --unit=function judges each function and prints its lines, with -n giving file line numbers:

$ ./sys1grep -n --unit=function -e "retries on network failure" src/net.js
src/net.js:3:async function fetchWithRetry(url) {
src/net.js:4:  for (let i = 0; i < 5; i++) {
src/net.js:5:    try { return await fetch(url); }
src/net.js:6:    catch { await sleep(2 * i  100); }
src/net.js:7:  }
src/net.js:8:}

A function runs from a funcname line to the line before the next one, as git grep -W finds it; lines before the first funcname line are one unit, and blank lines at a function's end are not printed. The funcname lines are git's: the diff= attribute in .gitattributes, then diff..xfuncname in git config. Without one, sys1grep has its own rule for JavaScript/TypeScript (a top-level function, class, or const / let / var bound to a function) and Python (def / class, nested ones too); a decorator line (@retry) starts the function it decorates. Otherwise it uses git's default: a line starting with a letter, _ or $. git's own builtin drivers (diff=python and so on) are not read, so a driver without an xfuncname falls back the same way:

$ printf '*.go diff=golang\n' >> .gitattributes
$ git config diff.golang.xfuncname '^(func|type)[[:space:]]'

-M defaults to 8000 characters here, as with -z; a longer function is judged on its first 8000. --chunk counts functions, -A/-B/-C and -c count lines, as with the sentence units. -z, -g and -o are refused.

From one function to another (--step-to)

A question about code often comes in two steps: where does --summarize hand the lines to the tool, and then, what happens when that fails. The "that" in the second step is what the first one found. Multi-step matching answers both in one run: the expression before --step-to finds the start, sys1grep walks the calls from it, and the expression after --step-to picks the functions it reaches that match. Jev judges the two expressions; no score decides which call to follow.

$ sys1grep -e "--summarize hands the lines to the tool" --step-to "what happens when the tool fails to start or answer" sys1grep.mjs

A meaning may start with a dash, as the first one does, on any search and for -e, -a, -v and -Q: a value that starts with - is taken as the meaning unless it is an option (an option holds no space).

Each path prints as a tree: the hop (the number of calls from the start), file:line and the function's name, : after an end and - after a function on the way, as grep marks a match and its context. A regex-only start and end send nothing, which makes a free reachability search. In Python 3.6's json, the functions that load reaches and that hold a raise:

$ cd /usr/lib64/python3.6/json && sys1grep -e '/^\sdef load\b/' --step-to '/\braise\b/' .py
sys1grep: walk: 1 at hop 0, 1 at hop 1, 3 at hop 2, 1 at hop 3, 1 at hop 4, 1 at hop 5; stopped: no new unit
0 __init__.py-274-load
  1 __init__.py:302:loads
    2 decoder.py:334:decode
      3 decoder.py:345:raw_decode
        4 scanner.py-65-scan_once
          5 scanner.py:28:_scan_once
  • The expressions. The expression before --step-to is the start's, as in any search; every -e / -a /
-v / -Q after it is the end's. --step-to MEANING and --step-to '/RE/' are --step-to -e ...; the file names follow. A step is a search, not a hop: one --step-to is two steps, whatever the hops between them.
  • The edges. Without --edges the unit is a function (--unit=function) and an edge is a call: name( in
a function's body, its comments, docstrings and strings left out, links to every function of that name. A name comes from the funcname line (the def under a decorator). --edges=FILE walks any other relation instead, one edge per line, FROM_FILE:LINEFROM_NAMETO_FILE:LINETO_NAME; a line stands for the unit that holds it. --reverse walks the edges backwards, callee to caller.
  • The walk. Breadth first from every start at once. A function's hop is its shortest distance from any
start, and each function is reached once, so a cycle ends. --hops=N, --hops=M..N and --hops=M.. pick the hops an end may be at (default 0..: a start that matches the end expression is an end at hop 0). A path that reaches no end is left out. stderr says how many functions each hop reached and why the walk stopped: no new function, or --hops.
  • The cost. Jev judges every function against the start expression in one batch, then the functions reached
within --hops against the end expression in another. --max-cost asks again before the second, counting the first. --dry-run answers every question 0, so it walks from every function the start expression could hold for and shows that bound.
  • Not with -z, -g, -o, -c, -l, -A/-B/-C, --rank, --summarize or --dedup. -p adds the
end expression's scores to each function judged against it; -q prints nothing and exits 0 when there is an end.

One line per template (--dedup)

Cost is proportional to the text sent, and machine-generated logs are mostly one skeleton with a different id or number in it. --dedup masks ids, hashes, numbers, dates and times, paths and URLs, groups lines by the result, sends one line per group and reuses its answer for the rest. What is sent is that line's original text, and every line is printed as itself:

--dedup takes auto, always or never; a bare --dedup is --dedup=always. **Default is never, for now** (#143): auto is held back until real-run stats say it should be the default. Every run still estimates locally, before anything is sent, what folding every kind (the best case) would save, and prints one stderr hint when that is worth at least twice the cost of --dedup's own question:

$ sys1grep -e "a request failed" app.log
sys1grep: 2040 units fold to at most 31 templates; --dedup=auto would save ~68 requests (~44k tokens)

--dedup=auto asks and folds only when that estimate says it pays; a file that barely repeats sends every line, asking nothing extra. --verbose / --dry-run print the decision and its numbers for every value.

$ ./sys1grep --dedup -n -e "a request failed" app.log
1:worker request 3fa9c1e27b failed: connection reset
2:worker request 88d0e41a5c failed: connection reset
4:worker request 0b7f2a9e13 failed: connection reset
3 of 6 lines matched; 4 sent to Jev (2 folded by --dedup, ~235 input tokens / ~$0.000010 saved, 21%) in 2 requests, 907 input tokens, ~$0.000038

Whether a value may be folded depends on the meaning: a number decides "disk usage is above 90%", a time decides "happened at night". So Jev is first asked, one small request per meaning, which kinds of value could change a match, and those kinds are kept apart (one of the two requests above). With -e "disk usage is above 90%" the same file keeps 95% and 12% apart and prints only 5:disk usage 95%.

Measured on real logs, with nothing kept: a 43,071-line system log folds into 550 templates (1.2% of the bytes), install.log to 21.1%, a Claude Code transcript (jsonl) only to 56.8%. It is for machine-generated logs; prose has no shared skeleton, and a meaning that reads a timestamp folds almost nothing. With -z or --unit=sentence-by-* the records or sentences fold instead of lines.

After a search, the summary line says what the folding saved: (2 folded by --dedup, ~235 input tokens / ~$0.000010 saved, 21%) above. It estimates what the folded lines would have cost as requests of their own and subtracts what the value questions cost. On a small file or on prose, the result can be negative.

To see how far your own logs fold before paying for a search, node scripts/dedup-measure.mjs FILE... counts lines, templates and the share of bytes sent, offline, with the same masks; --keep=num,time shows a meaning that keeps those kinds apart.

The request's other lines are each line's context (#9), and --dedup changes them, so a line near the threshold can be judged differently than in a full pass.

A summary instead of the lines (--summarize)

When many lines match, what you want is often the gist: what they say about the meaning you searched for. --summarize pipes what sys1grep would print to claude -p --model haiku (no tools, no settings, no CLAUDE.md), asks it to summarize the lines as they bear on the meanings, and prints its answer instead of the lines.

$ sys1grep -r -n --summarize -e "the API key is read from a file" .
The key is read in sys1grep.mjs:22 from ~/.config/sys1grep/.env, never from ./.env (sys1grep.mjs:18), ...

The expensive model reads only what Jev kept. Asking why ./.env is no longer read of this repository's git log (168 commits), Claude's input fell from 22,059 tokens to 1,356, the total cost with Jev's from $0.094 to $0.013, with the same answer (#69). It pays when the answer sits in a few lines.

  • The matching lines leave the machine a second time, to Anthropic (or whatever TOOL talks to).
  • -n, -A/-B/-C, -p and file names go in as they would print; colors never do. No match runs nothing (exit 1).
  • SYS1GREP_SUMMARIZER picks the TOOL of a bare --summarize, SYS1GREP_SUMMARIZER_MODEL its model.
  • -q, -l and -c print no lines, so they cannot be combined with it.
  • In SYS1GREP_OPTS it summarizes every search, so every match goes to TOOL's provider too;
--no-summarize turns it off for one search.
  • The answer is plain text: an LLM that is not told writes Markdown, noise on a terminal, so plain is asked for.
--format=markdown or =html asks for those instead (… --format=html … > summary.html). The answer prints as it comes, unchecked. In SYS1GREP_OPTS it is a standing preference, ignored without --summarize or --rank.
  • With --dedup, the TOOL gets what Jev got: each template's representative once, marked (×N like it),
not every line its answer was reused for.
  • More than 200 KB (about 50k tokens) is not sent at all: exit 2, with the size, before the TOOL is paid.
It is never cut short, since a summary of the first part would read as a summary of all of it.

--summarize-prompt=TEXT adds your own instruction after the fixed one (how long, what to focus on):

$ sys1grep -r -n --summarize --summarize-prompt="3 lines or fewer, just which file to fix" \
    -e "the API key is read from a file" .

It needs --summarize; empty TEXT is the same as leaving it out. It comes after the --format sentence, so TEXT can override the format.

Other TOOLs:

$ sys1grep --summarize=llm -e "..." FILE            # Simon Willison's llm, tools off (no -T given)
$ sys1grep --summarize=pi -e "..." FILE             # pi --print --no-tools --no-session ...
$ SYS1GREP_SUMMARIZER_MODEL=qwen3.5:9b sys1grep --summarize=ollama -e "..." FILE
$ SYS1GREP_SUMMARIZER_MODEL=some-id sys1grep --summarize=lmstudio -e "..." FILE   # model id from GET /v1/models
$ SYS1GREP_SUMMARIZER_MODEL=some-id sys1grep --summarize=http://localhost:8080/v1 -e "..." FILE  # llama.cpp, vLLM, LocalAI, a gateway

ollama and lmstudio talk to a local OpenAI-compatible server (POST /v1/chat/completions) by fetch, no CLI: the matching lines never leave the machine a second time. A plain http(s):// URL is any other OpenAI-compatible server. All three need SYS1GREP_SUMMARIZER_MODEL: none has a default model. SYS1GREP_SUMMARIZER_API_KEY goes as Authorization: Bearer to a URL TOOL only (never SYS1GREP_API_KEY, which is Jev's). OLLAMA_HOST moves ollama's host, as it does for the ollama CLI itself.

Best first (--rank)

Matches print in file order, as grep prints them. With many, --rank prints the results best first, each under a numbered header. A result is a match with its -A/-B/-C lines; matches whose context touches are one result.

$ sys1grep -n -C1 --rank -e "the refund was refused" tickets/
  1. tickets/b.txt
tickets/b.txt-2-Order #1234, placed 2026-08-01. tickets/b.txt:3:Your refund was declined: the order is older than 30 days. tickets/b.txt-4-Please contact support.
  1. tickets/a.txt
tickets/a.txt-11-Order #88 arrived. tickets/a.txt:12:Refund request received. tickets/a.txt-13-We will check it.
  • --rank is --rank=jev: after the search, Jev is asked of each result, lines and context together, whether it
is relevant to the meanings that are not negated. One more question per result, --chunk results to a request; --dry-run / -i show an upper bound, since which results there are is known only after the search.
  • --rank=match sorts by each result's highest match probability, with no request. It is the answer to a yes/no
question on one line, so clear matches sit close together, and the context is not read.
  • -p puts the score on the header (1. [0.96] tickets/b.txt). -l lists the files by their best result.
  • --format=markdown writes a ## 1. tickets/b.txt heading and a fenced block per result; --format=html one
self-contained HTML document from a template (below), light and dark, with a relevance bar per result. The lines are as they print, escaped and uncolored.
  • --summarize gets the results in ranked order; with --dedup a representative is one result.
  • It needs a
... (README truncated for length)
Chat with me