release.yml TOOLS + BIN cases, VS Code build/test tasks, README table, plan.md milestone 5, tool-parity build-order status. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011KikHkfCiC3yELbsMN8fT9
14 KiB
Tool parity: Google's agent vs Claude Code
Goal: let Google's agent CLI do everything Claude Code can, by building the missing capabilities as small CLI tools in this repo. The agent calls them through its shell tool, so every tool follows the house rules: text in, verifiable artifacts out, single static Go binary, self-explanatory output.
Which Google agent? Legacy gemini-cli is auth-dead for personal
accounts (verified again 2026-08-05, see infra-Doc
hosts/brasse-linux01.md). The real target is agy (Antigravity
CLI) — live-tested in section 2, and its toolset differs from the
old Gemini CLI docs. Section 1's table is kept for reference since
Gemini CLI still exists in API-key mode.
A design rule that fell out of the live test: a dedicated binary
beats ad-hoc shell because of approval prefixes. agy (like Claude
Code) allowlists commands by prefix — notifyr … can be approved once
and forever, while every hand-rolled for i in $(seq …); do curl …
loop is a unique string that needs fresh human approval. Small stable
CLIs are therefore not just convenience: they are what makes
unattended agent operation possible at all.
Sources: Gemini CLI tools reference (https://geminicli.com/docs/reference/tools/),
Claude Code's toolset as of 2026-08, live probing of agy 1.1.9.
1. Already at parity — nothing to build
| Capability | Claude Code | Gemini CLI |
|---|---|---|
| Read/write/edit files | Read / Write / Edit |
read_file / write_file / replace |
| Find files / search text / list dirs | Glob / Grep |
glob / grep_search / list_directory, plus read_many_files |
| Shell, incl. background processes | Bash (+ background tasks) |
run_shell_command (+ background processes) |
| Web | WebFetch / WebSearch |
web_fetch / google_web_search |
| Ask the user a structured question | AskUserQuestion |
ask_user |
| Plan mode | EnterPlanMode / ExitPlanMode |
enter_plan_mode / exit_plan_mode |
| Skills / slash commands | Skill (.claude/skills) |
activate_skill (.gemini/skills) |
| Persistent memory | file-based memory dir | save_memory (simpler, but exists) |
| Todo/task tracking | TaskCreate/TaskUpdate/… |
write_todos, tracker_* (experimental) |
| MCP servers + resources | MCP tools, ListMcpResources/ReadMcpResource |
MCP tools, list_mcp_resources/read_mcp_resource |
| Subagents | Agent (background, custom types) |
subagents (experimental) — weaker, see fanout below |
Not worth replicating (harness-internal to Claude Code, no value as a
CLI): ToolSearch, EndConversation, ReportFindings,
ShareOnboardingGuide, DesignSync, remote cloud execution.
2. Live test 2026-08-05: agy (Antigravity CLI 1.1.9)
Tested interactively in a tmux session (Google AI Pro account, model Gemini 3.6 Flash). Its 19 built-in tools, self-enumerated:
ask_permission, ask_question, define_subagent, generate_image,
grep_search, invoke_subagent, list_dir, list_permissions,
manage_subagents, manage_task, multi_replace_file_content,
read_url_content, replace_file_content, run_command, schedule,
search_web, send_message, view_file, write_to_file.
What this changes vs the old Gemini CLI picture:
- agy has real subagents (
invoke_subagent/define_subagent/manage_subagents+send_message) and background-task management (manage_task). →fanoutdemoted further; probably never needed. - agy has
schedule— one-shot timer or cron expression that wakes the agent with a prompt (same idea as Claude'sScheduleWakeup). Confirmed limits, from its schema: it cannot run commands itself, it is in-memory and dies with the session, and it cannot reach the phone. →cronr(persistent systemd timers) andnotifyrare still needed;schedulecomplements them within a session. - agy has
generate_image— a reverse gap: Claude Code has no native image generation. Nothing to build; just worth knowing. - No MCP-resource tools, no memory tool and no glob in its toolset (grep/list_dir cover finding files).
Behavior tests run in a scratch arena:
| Test | Result |
|---|---|
| Enumerate tools | Clean list of 19 (above) |
| Edit a text file | Worked, auto-approved in trusted folder |
Edit a Jupyter cell, keep .ipynb valid |
Passed — it wrote a python3 -c json script rather than text-replacing. Notebook stayed valid. Cost: a per-command approval each time → nbcell demoted to nice-to-have (stable prefix + no ad-hoc python). |
| Wait for a file to appear | Worked via a hand-rolled for … sleep 1 shell loop — a unique command string needing fresh approval. → exactly the waitfor case. |
| Asked agy which CLI tools it wants for the homelab | Its list: ntfy client, tea (Gitea CLI), skopeo/crane (registry), ofelia/cron daemon, ctop-style fleet status — near-1:1 with section 4, and it independently made the approval-prefix argument. |
Buy before build: agy's suggestions overlap with off-the-shelf
tools. Evaluate first: tea (official Gitea CLI — but it does not read
the Pi5's zst action logs and has no wait-quiet, which stay
giteactl's reason to exist, possibly as a thin layer on top of
tea), skopeo/crane (cover most of reghelper — remaining value
is prune plans and size summaries), ctop (interactive TUI, not
agent-friendly output — fleet still wins for agents).
3. Gaps → tools to build
Ordered by expected value. Each becomes its own folder + binary + README, per repo convention.
3.1 cronr — scheduled/recurring runs (Claude: CronCreate/CronList/CronDelete, /loop)
agy's built-in schedule dies with the session (see section 2).
cronr manages systemd user timers so an agent can create
recurring or one-shot jobs that survive session exit and reboot
(including "run this prompt every morning" via agy -p …).
cronr add nightly-ci-check --schedule "*-*-* 07:00" --cmd 'agy -p "check CI status and notify"'
cronr add once-reboot-check --at "2026-08-06 03:00" --cmd '…' # one-shot
cronr list # name, schedule, next run, last result
cronr logs nightly-ci-check # journalctl for the unit
cronr rm nightly-ci-check
Output prints the generated unit files so the result is verifiable. No daemon of its own — systemd does the running.
3.2 waitfor — block until a condition holds (Claude: Monitor)
Turns "poll every N seconds" into a single blocking tool call, so the agent doesn't burn turns polling.
waitfor --cmd "curl -sf https://gitea.brasse-pc.eu/api/healthz" --interval 30s --timeout 20m
waitfor --cmd "ssh pi5 docker ps --format '{{.Names}}'" --matches 'gitea' --timeout 10m
waitfor … --then 'notifyr send --msg "gitea is back up"'
Exit 0 = condition met, exit 3 = timeout; last output is printed either
way. --then runs a command on success (composes with notifyr).
3.3 notifyr — push notifications, send and read (Claude: PushNotification)
The ntfy server already runs on the Pi5 and is the house-wide
notification bus (topics like Info, pi5-server-fel, ci-fel — see
infra-Doc services/observability.md). notifyr is a thin client so
every agent uses it the same way:
notifyr send --topic ci-fel --title "Build failed" --msg "agent-tools arm64 test: FAIL" --priority high
notifyr read --topic pi5-server-fel --since 2h # poll mode: what has alerted lately?
read (ntfy's ?poll=1&since=…) is the underrated half: it lets an
agent check what the infra has been complaining about before/after a
change. Config (~/.config/notifyr/config.json): server URL + token.
3.4 pagepub — publish an HTML/Markdown report to a URL (Claude: Artifact)
Claude Code can publish reports as web pages; Gemini cannot. pagepub
rsyncs a file to a static-file host on the Pi5 (nginx container behind
NPM, e.g. pages.brasse-pc.eu) and prints the stable URL.
pagepub publish report.html --slug ci-report → https://pages.brasse-pc.eu/ci-report/
pagepub publish notes.md --slug pi5-audit # .md rendered to HTML with built-in template
pagepub list | rm <slug>
Requires the static host to exist first (small infra task; goes in
infra-Doc + NPM proxy host via npmctl).
3.5 nbcell — Jupyter notebook editing (Claude: NotebookEdit)
.ipynb is JSON that's miserable to edit via replace. nbcell
exposes cells as text:
nbcell list nb.ipynb # index, type, first line, exec count
nbcell show nb.ipynb 3 # cell source (and outputs with --outputs)
nbcell edit nb.ipynb 3 --from-file cell.py
nbcell add nb.ipynb --at 4 --type code --from-file new.py
nbcell rm nb.ipynb 7
3.6 wtreectl — disposable git worktrees (Claude: worktree isolation for agents)
Claude Code can give each subagent an isolated git worktree. wtreectl
does the same for any agent:
wtreectl new [--branch dev/foo] # prints the new worktree path
wtreectl list
wtreectl clean # removes worktrees with no changes
Lets two agent sessions work in the same repo without trampling each other.
3.7 fanout — parallel subagent orchestration (Claude: Workflow, Agent) — roadmap, not agreed
Runs N prompts as parallel headless agent processes (gemini -p /
claude -p) with a concurrency cap, collecting each result as JSON in
an output dir. A poor man's Workflow:
fanout run jobs.json --max 3 --out results/
Heavier than the other tools and overlaps with agent-helm's territory — park until there's a concrete need.
4. Suggested tools from infra history
Grounded in what the agent has already been doing per infra-Doc
(maintenance-and-gaps.md, services/source-control-and-deploy.md,
services/observability.md, the per-service "operational quick-ref"
blocks). These help any agent (Claude or Gemini) administer the
fleet, and they respect the read-only sudo policy
(ssh/claude-sudo-policy.md): everything below is read-or-notify;
mutations still go through Björn's supervised tmux flow.
4.1 giteactl — Gitea repos, Actions runs and CI logs
The biggest recurring friction. Today: CI status is polled ad hoc, and
logs for private repos are only readable as zst files under
/srv/storage1/gitea/actions_log/… on the Pi5. Wraps the Gitea REST +
Actions API:
giteactl runs <repo> [--limit 5] # status, branch, duration
giteactl log <repo> <run> [--job N] # fetches + decompresses the zst log
giteactl wait <repo> [--timeout 30m] # block until latest run finishes; exit 0 = green
giteactl wait-quiet [--max-active 1] # block until ≤N heavy builds are running
giteactl release <repo> [<tag>] # rolling-release assets + checksums
wait-quiet encodes the hard-learned rule "serialize pushes — >2 heavy
builds take the Pi5 down": giteactl wait-quiet && git push.
Composes with waitfor/notifyr.
4.2 fleet — one-shot health snapshot of the whole homelab
Every service doc ends with the same hand-rolled loop over
ssh pi5-claude sudo claude-docker ps/inspect …. fleet does that
loop once, properly:
fleet status # containers (state, image, restarts), disk/mergerfs fill %, failed systemd units
fleet status --host brasse-linux01
fleet checks # Uptime Kuma monitor states + last ntfy alerts (via notifyr read)
Read-only by construction (claude-docker wrapper + sudo allowlist), so it needs no new permissions. Output is a stable text table an agent can diff between runs.
4.3 envaudit — compose ↔ .env key auditor
Grounded in a real incident (${STORAGE1} undefined in /srv/.env →
bad mount, rollback) and in maintenance-and-gaps' inline-secret
findings:
envaudit check /srv/dockge-staks --env /srv/.env
→ UNDEFINED ${STORAGE1} used by media-stack/compose.yaml
→ INLINE LDAP_ADMIN_PASSWORD hardcoded in openldap/compose.yaml (should live in .env)
→ UNUSED OLD_API_KEY defined but referenced nowhere
Pure text analysis of compose files — safe to run anywhere, catches the two failure classes that have actually happened.
4.4 reghelper — docker-registry catalog & hygiene
The LAN registry (192.168.0.19:5000) has no UI, no auth and no
cleanup story, and the backup plan explicitly wants it slimmed:
reghelper ls # catalog + tags + image sizes
reghelper tags <image>
reghelper prune-plan --keep 2 # prints the delete+GC commands (does NOT run them)
prune-plan deliberately only prints the mutation commands for the
supervised tmux flow — same pattern as the sudo policy.
4.5 Backup status reader — once Backrest/restic exists
future-plans.md has the whole Backrest+restic design chosen but
unbuilt. When it lands, a fleet backups subcommand (last snapshot age
per source, repo size, last check result) closes the loop — an agent
can then verify backups instead of trusting them. Not a separate
tool; park under fleet.
Cross-reference
/home/brasse/repos/dify-agent-tools/ already has a specced-but-unbuilt
set of FastAPI tools (file/image sorting, face recognition, ST-card
export…) for the Dify platform. Different runtime (HTTP tools vs CLI
binaries), same philosophy — don't duplicate those here.
5. Suggested build order
- ✅
notifyr— smallest, everything else composes with it, ntfy already runs. (built 2026-08-07) - ✅
giteactl— removes the biggest daily friction (CI logs + build serialization). (built 2026-08-07; private repos need an API token in the config) - ✅
waitfor+cronr— turns both agents into unattended operators. (built 2026-08-07) fleet— replaces the hand-rolled health loops in every runbook.envaudit,reghelper,pagepub,nbcell,wtreectl— as needed.fanout— only if a concrete multi-agent need shows up.
Also built 2026-08-07 (outside this list, Björns direct request):
svgc (svg-maker/) — SVG graphics from text, displayed in
agent-helm via helmd share. Agy integration: the agent-tools
skill + a tools section in ~/.gemini/config/AGENTS.md.