Super-agent
The super-agent is a mode, not a persona name. When you run apx exec without targeting a
specific project agent, APX activates this mode: the daemon-level, tool-using loop that can reach
every registered tool — file operations, MCP calls, Telegram, shell, tasks, sessions, and more.
The name the assistant uses when talking to you comes from ~/.apx/identity.json (default
“APX”). References to “super-agent” are an internal technical term for the mode. You configure
it with super_agent.* keys in ~/.apx/config.json, and routine kinds may still say
super_agent — but users never see that label.
Enabling it
Section titled “Enabling it”The mode is disabled by default. Set both keys before you first call it:
apx config set --global super_agent.enabled trueapx config set --global super_agent.model "anthropic:claude-sonnet-4-5"Once enabled, every apx exec call without -a <agent> goes through this mode.
apx exec — the one-shot entry point
Section titled “apx exec — the one-shot entry point”apx exec "<prompt>" # super-agent mode (default)apx exec -- "<prompt>" # same; -- avoids flag parsing issuesapx exec -a reviewer "<prompt>" # named project agent insteadapx exec "<prompt>" --model gpt-5 # override the configured model for this callapx exec "<prompt>" --max-tokens 2000 # cap output tokensapx exec "<prompt>" --project my-app # pin to a specific project context$ apx exec 'list my open tasks across all projects' Roby (super-agent) · anthropic:claude-sonnet-4-5 → tool tasks.list { project: "First Project", state: "open" } → tool tasks.list { project: "Acme Store", state: "open" } → tool tasks.list { project: "Data Pipeline", state: "open" } You have 4 open tasks: • First Project — Wire Stripe webhooks before merge (due 06-16) • Acme Store — Retry-guard the flaky cart-total test (due 06-18) • Acme Store — Document the pairing flow for the web panel • Data Pipeline — Backfill ingestion CLI flag (due 06-20) # 612 tok in / 198 out · 2.4s
When the super-agent mode is active the daemon resolves the assistant’s name from
~/.apx/identity.json and displays it in status lines, TUI headers, and Telegram replies —
never the string “super-agent”.
Reaching it from another agent (A2A)
Section titled “Reaching it from another agent (A2A)”The configured identity is also a first-class A2A address. An agent can contact Roby as naturally as any project agent:
apx send crypto-analyst roby "BTC alert: decide whether the owner needs this on Telegram" --deliverAPX finds the sender’s project, files every identity alias under the stable super_agent peer,
and runs the real super-agent loop: identity, relevant memory, project context, permission policy,
and tools. This is different from apx exec, which represents an owner-authored turn. If the
sender slug exists in several projects, add --project <name>.
Real project agents still win over aliases. If a project intentionally defines an agent whose
slug is apx or whose name matches the configured identity, address selection chooses that agent
inside its project.
The tool loop
Section titled “The tool loop”Under the hood runSuperAgent() builds a system prompt (identity block + project inventory +
memory block + active-threads block), selects a tool schema subset appropriate for the channel,
then calls the engine via the standard runAgent() loop. The loop continues until the model stops
calling tools or a signal aborts it.
Tool registry
Section titled “Tool registry”The full registry has around 30 native tools plus registry-bridged tools (browser, fetch, search, glob, grep). The daemon picks a subset based on the channel to stay within cheap-tier TPM budgets:
| Channel | Tool set | Rationale |
|---|---|---|
api / apx exec | Full | Deliberate, user-picked model |
routine | Full | Scheduled, autonomous, user-picked model |
web / code | Full | Long-form workspace |
telegram | Core (~700 tokens) | Cheap tiers, snappy replies |
desktop / deck | Core | Same |
The model can always expand its surface by calling load_skill.
Selected native tools
Section titled “Selected native tools”| Tool | What it does |
|---|---|
create_task / list_tasks / get_task | Add, list and read per-project TODO items |
update_task / complete_task / comment_task | Edit one, move or close it, annotate its thread |
call_agent | Hand work to a named project agent, in-process. It runs with its own tools, so it can do the work rather than describe it, and the exchange is filed as an agent-to-agent conversation you can read and continue |
send_to_agent | Message another agent — a project agent, the super-agent, or a coding runtime — and get their answer. With background: true the work is left running and the sender is woken when it lands, so a long exchange no longer holds a turn. See Work left running |
call_mcp | Call any registered MCP server tool |
call_runtime | Delegate to an external runtime (Claude Code / Codex …), synchronously or in the background |
run_subagent | Spawn an isolated sub-agent (fresh context, same tools minus user-interaction ones) that works a self-contained task to completion and returns its result — one nesting level max |
run_shell | Execute a shell command (gated by permission mode) |
send_telegram | Send a message via the Telegram plugin |
send_file | Hand you a file from disk — a screenshot, a generated image. See Sending you files |
remember | Write a durable fact to ~/.apx/memory.md |
search_sessions | Retrieve past session transcripts |
set_identity | Update ~/.apx/identity.json fields |
Sending you files
Section titled “Sending you files”The super-agent can hand you a file with send_file: a browser screenshot, an
image it generated, a render. The file arrives with the reply: in the web
chat it shows inline with a download button, and on Telegram it goes out as a
photo (or as a document for other kinds). Taking a screenshot does not send it
by itself; the agent has to pass the saved path to send_file.
Only the kinds you enable are sent. Images are on by default, and everything
else is off until you turn it on in ~/.apx/config.json:
{ "file_delivery": { "kinds": { "image": true, "video": false, "audio": false, "document": false }, "max_mb": 20 }}document covers PDF, text, Markdown, CSV, JSON, Office files and .zip.
Some things are never sent, whatever the config says: executables and scripts
(.exe, .msi, .bat, .sh, .apk, .dmg, …), and anything under
~/.ssh, ~/.aws, ~/.gnupg or ~/.apx (except ~/.apx/media, where sent
files are archived). An image whose bytes don’t match its extension is refused
too, so renaming a file doesn’t get around the switch.
Work left running
Section titled “Work left running”A peer’s turn is a full tool loop. It can take minutes, and send_to_agent waits
for it by default — which is right only when the answer is the next thing the
sender needs.
background: true changes that. The tool returns immediately with a job_id,
the peer works on its own, and when it answers the sender is woken: the reply
arrives as a new message on the same agent-to-agent thread, starting a fresh turn
with the result in hand. Pass wake_me: false for a pure notification nobody
needs an answer to.
The wake-up is an ordinary a2a message in the other direction, so the request and its answer live in one thread — there is nowhere new to look.
What the woken sender writes is filed as its own note, not as a reply: the peer already answered and is waiting on nothing. Filing it as a reply used to put “received, thanks” in the peer’s inbox, the peer confirmed back, and one answer turned into a round of acknowledgements, each a full tool loop.
Rules for handing work on
Section titled “Rules for handing work on”- Asking twice does not start it twice. While a hand-off to an agent is still running, calling
that agent again (
call_agentorsend_to_agent) returns the running job — what it is doing and for how long — instead of opening a new turn.check_jobsreads the same status without touching anyone. A deliberate correction goes through withfollowup: true. - Nobody waits from inside an a2a thread. An agent answering another agent hands work on in the background, never by waiting on a third one; heavy work (a render, a production) runs as its own background job, not inside the conversation.
- No question cards in a2a. Nobody there can answer one: if an agent needs the owner, the super-agent asks on its own channel.
- The chat shows each hand-off as Waiting / Waited or Background, and the composer bar of the super-agent’s chat lists every agent working for it, with the chain below each one.
What the sender keeps, and what it does not
Section titled “What the sender keeps, and what it does not”A woken turn is a new turn. It gets the reply, a recap of what was asked, and the thread’s own history — not the working context the sender had when it handed the job off. Anything needed to act on the answer belongs in the message itself.
A job that produced no answer wakes the sender too, and says so plainly: the peer failed, the job ran past its budget, or the daemon restarted while it ran and the work is unrecoverable. The sender is told not to report such a job as done.
Limits
Section titled “Limits”| Limit | Value | Why |
|---|---|---|
| Jobs open per agent | 3 | Bounds how WIDE the fan-out goes |
| Hand-off depth | 3 | Bounds how DEEP it goes; counted across turns, wake-ups, call_agent and apx send |
| Tool steps per peer turn | 20 (super_agent.a2a_max_iters) | A peer that runs out answers with what it did and what is left; the sender decides whether to continue |
| Job lifetime | 1 hour by default | Past it the job is cut off and the sender told |
Both limits come back as a message the agent can act on, never as a crash.
A turn nobody is watching — an a2a reply, a routine — also stops on a spent
plan. When a provider answers that the account’s usage limit is reached (not a
passing rate limit), that model cools down for an hour (super_agent.quota_cooldown_min),
the unwatched turn ends instead of spending the rest of the fallback chain, and
the owner gets one Telegram line saying so. A chat with a person in it still
falls through to the next model. So does a model you chose — pinned to the
agent, set on the routine, or the super-agent’s own model — when its fallback
switch is on: that switch is your instruction to carry on down the router. The
stop is for the router’s own account running dry under agents that inherit it.
Seeing what is running
Section titled “Seeing what is running”curl -s localhost:7430/api/jobs?open=1 -H "Authorization: Bearer $TOKEN"GET /api/jobs lists every job — ?open=1 for what is running now, plus
?project_id, ?from and ?status. GET /api/jobs/:id returns one. The panel
reads the same route, and the live events feed carries a background_job frame
on start and end.
Jobs are read-only over HTTP on purpose: one ends when its work ends, its deadline passes, or the daemon holding it dies. To stop the work, abort the turn the job opened.
Cross-channel brokering
Section titled “Cross-channel brokering”The super-agent mode is the common handler behind every surface: Telegram messages, Desktop window
prompts, apx exec calls, web chat, and routine runs all converge on runSuperAgent(). Each
surface passes a channel identifier (telegram, desktop, web, routine, api, …) which
drives:
- Which tool subset is visible to the model.
- Whether an “active threads” block (recent turns on other channels) is injected.
- Which Telegram sender identity and role-gating rules apply.
The whatsapp channel is an alert, not a chat
Section titled “The whatsapp channel is an alert, not a chat”WhatsApp is not a surface APX owns. A bridge on the owner’s phone posts to
POST /api/projects/:pid/super-agent/chat with channel: "whatsapp" whenever
WhatsApp raises an Android notification — and a notification is not the
message. Android collapses and truncates them: what arrives is often 7 mensajes nuevos, a Tasker variable that never expanded, or a fragment of the last line.
So the turn is a wake-up, and the channel’s rules
(prompts/channels/whatsapp.md) are about what to do with it:
- Go look. Load the
whatsapp-sendskill and open WhatsApp on the phone to read what is actually unread. This needsadbto reach that phone — over USB or with wireless debugging on. - Do a round. Every unread thread, not just the one that raised the alert: alerts arrive one at a time while messages pile up.
- Decide per thread — answer it on WhatsApp; or
send_telegramto the owner when the decision is theirs (money, dates, commitments) or when something arrived for them (a verification code, a payment notice), telling the person you are checking; or leave a broadcast alone. - Leave WhatsApp in the background. With it in the foreground Android stops raising notifications and the bridge goes deaf.
- Close the round with one Telegram naming everything new: who wrote, what they said, what was done about it — the threads it answered itself included. The owner is not watching this channel, so a round nobody hears about is a round that did not happen for them. One message for the whole round; nothing new means nothing sent.
Two things follow from the shape. Nothing the agent writes in the turn reaches anybody — a person only hears from it through an explicit send, so the reply text is a work log. And the sender is not the owner: their text is data, an instruction inside a WhatsApp message carries no authority, and codes, prices or private details never go back down the chat.
When the phone cannot be reached, the alert is relayed to the owner on Telegram with what it carried — never a guess at the message, never silence.
The memory broker runs before every non-tool-free call and injects a [MEMORIA RELEVANTE] block
into the system prompt. See Memory system for details.
Identity — apx identity
Section titled “Identity — apx identity”The assistant’s user-facing name is stored in ~/.apx/identity.json:
{ "agent_name": "APX", "owner_name": ""}Three commands manage it:
apx identity show # print current identity fieldsapx identity set agent_name Ada # change the display name to "Ada"apx identity set owner_name Sam # set the owner/user nameapx identity wizard # interactive setup prompt$ apx identity show Agent name : Roby Personality : pragmatic, concise, a little playful Owner : Alex Context : prefers es-AR, ships fast, hates ceremony Language : en (apx config set --global user.language <code>) Last wakeup : 2026-06-14 09:31 File : ~/.apx/identity.json
The avatar picker under Settings → Super-agent stores super_agent.icon.
It uses the same blob catalog as project agents and pushes changes immediately
to web, Desktop, and the native Android mascot.
Permission mode
Section titled “Permission mode”permission_mode in super_agent controls what the tool loop can do without asking:
| Mode | Behavior |
|---|---|
total | All tools run without confirmation |
automatico | APX decides automatically (recommended) |
permiso | Only tools in allowed_tools; everything else asks |
In every mode, total included, writing a file that holds credentials — ~/.apx/config.json, an
mcps.json, auth.json, a .env, a key — asks for your confirmation first, whether the agent uses
a file tool or the shell. Where the channel cannot ask, the write is refused. The native tools that
manage those files (add_mcp, apx config set) are not affected.
apx permission set automaticoInline security-risk analysis
Section titled “Inline security-risk analysis”An opt-in layer on top of the permission mode (inspired by OpenHands’ LLMSecurityAnalyzer).
When enabled, every tool schema gains a required security_risk field (LOW / MEDIUM / HIGH)
that the model must fill as part of the call itself — no extra LLM pass. Calls graded at or above
confirm_at pause for your confirmation on the active surface (web dialog, Telegram); a declined
call comes back to the model as an error observation so it can re-plan.
"super_agent": { "security_risk": { "enabled": true, "confirm_at": "HIGH", "confirm_unknown": true }}confirm_at— minimum grade that pauses (LOWconfirms everything,HIGHonly destructive/outward actions).confirm_unknown— also pause when the model didn’t grade the call (weak models sometimes omit the field).
It’s not a duplicate of the permission mode — it gates by a different axis. Permission mode decides by tool identity (is this tool dangerous / on the allowlist); the risk analyzer decides by the model’s judgment of this specific action’s severity. So it adds a guard the permission mode can’t:
- In
automaticothe risk grade replaces the static dangerous-tool confirmation. - In
permisothe allowlist still applies and the risk gate composes on top. - In
totalit becomes a safety floor: everything runs free except aHIGH-graded action, which still stops for confirmation even under full trust. That’s the point — a catastrophic action gets caught no matter how much you trust the agent.
Goal-completion judge
Section titled “Goal-completion judge”Verification loop (OpenHands critic pattern): an LLM judge scores the probability that the
original request is fully satisfied. Below success_threshold the agent receives an internal
verification note (what looks unmet, per the judge) and continues — up to max_iterations rounds.
An unusable judge (engine down, unparseable reply) accepts the result rather than blocking it.
It covers two different ways a turn ends early:
enabled(opt-in) — a completion-contract turn (coding surfaces) declared itself done, and the judge checks that claim.continue_unfinished(on by default) — a conversational turn (Telegram, web chat) stopped calling tools. That is also what a model does when it merely announces its next step (“ahora genero el SRT en inglés”) and writes no call: the task stops halfway and waits to be poked. The judge is the poke. A turn that ran no tools is never judged (chat has nothing to finish), nor is one that ended by asking you something — that one is waiting, not unfinished. “Asking you something” covers both ways a turn can ask: theask_questionstool, and the reserved closing step that runs out of tool budget and offers to keep going. Continuing that one answers over your head, and since the next round hits the same wall it just repeats the same recap and the same question.
Project agents go through it too, not only the super-agent — through the
continue_unfinished half only, since they never take the completion-contract path. It applies on
every channel a project agent answers on, not just the watched ones: announcing a next step
instead of taking it is the model’s failure, not the surface’s. The two exclusions above are what
keep the cost down — after them the only turns judged are the ones that ran real tools and then
went quiet without asking anything, which is one extra model call for a turn that would otherwise
have sat there until you typed “seguí”.
"super_agent": { "judge": { "enabled": true, "continue_unfinished": true, "success_threshold": 0.6, "max_iterations": 2, "model": "" }}model empty → the judge runs on super_agent.model, and failing that on the model the turn
itself ran on — so a project agent is still judged on an install that never configured a
super-agent. Set a cheap model to keep verification overhead low. Verdicts surface as judge_verdict events, on the result as result.judge, and on
the Telegram message that closed the turn — so a turn that ran long can be traced back to the
verdict that kept it going.
max_iterations caps how many times the judge may send a turn back — it does not multiply
what the turn may spend. The rounds share the surface’s tool budget with the run that came
before them: each gets what is left, and when too little is left to be worth a round (one action
step plus the reserved closing one) the loop stops there and emits judge_budget_exhausted rather
than starting it. So one message on the web chat costs at most web_max_iters, not
(1 + max_iterations) × web_max_iters.
Per-surface tool budget
Section titled “Per-surface tool budget”How many tool steps one turn may take depends on whether you can watch it work — and it is keyed
on the surface, not on who is answering. A project agent in the web chat gets the same budget
Roby does; a watched chat is a watched chat. (All four keys live under super_agent because that
is where every tool-loop budget lives, including the routine one that governs project agents.)
| Surface | Budget | Why |
|---|---|---|
| Telegram | telegram_max_iters (1000) | Each real tool start is visible and Telegram keeps the typing indicator alive, so the turn normally runs until work is done. |
| Web chat + sidebar | web_max_iters (1000) | Every call renders live and you’re one click from stopping it — you are the guardrail, so the turn runs until the work is done. |
apx exec | cli_max_iters (40) | Room to hand work to a project’s agent and check it; bounded, because scripts and other agents call it too. From apx exec a hand-off to another agent waits for the answer instead of running in the background. |
| Routines (non-Telegram) | routine_max_iters (1000) | Nobody is waiting mid-run to answer “want me to keep going?”. |
| Group rooms | group_max_iters (50) | The one surface where a single message fans out — see below. |
| Coding surfaces | Explicit, plus the completion contract | They stop on finish, not on a count. |
Each number is the budget for a turn, not for one pass of the tool loop: if the completion judge sends the turn back to finish something, those rounds spend what is left of the same number rather than being handed it again.
The three 1000 ceilings are runaway backstops, not normal stopping points: the loop already ends
the moment the model stops calling tools, and stuck detection aborts the repeat-loops. Set any of
these to 0 for the built-in default.
A group room is watched and stoppable like the web chat, so by the surface rule it would take the same ceiling — but one line you write can cascade into up to ten agent replies, each a full tool loop. Giving every speaker the 1:1 ceiling would quietly make one message worth ten of them. So a speaker gets a slice instead: room for genuine multi-step work, while the ten of them together still cost no more than one turn in your 1:1 chat with any of them. Same principle as the judge rounds above — the budget belongs to your turn, and everything it fans out into fits inside it. A speaker that runs out closes with what it got done and the room carries on; solo work that needs hundreds of steps belongs in that agent’s own chat, which has the ceiling for it.
"super_agent": { "telegram_max_iters": 0, "web_max_iters": 0, "routine_max_iters": 0, "group_max_iters": 0}When a budget is exhausted, the last step is reserved for a tool-free, model-authored closing message: what it did, what’s left, and whether to keep going. With the run-to-completion ceilings, that is a rare safety floor rather than the way normal work ends.
Stuck detection
Section titled “Stuck detection”The tool loop watches for two loop shapes: the same call returning the same result
action_repeat times, and the same call erroring error_repeat times in a row. On the first
detection the model gets an in-band nudge (“you appear stuck — change approach or explain the
blocker”); if it keeps looping, the turn is closed early with a model-authored wrap-up instead of
burning the rest of the iteration budget. On by default:
"super_agent": { "stuck_detection": { "enabled": true, "action_repeat": 4, "error_repeat": 3 }}Model fallback
Section titled “Model fallback”If the primary model is unavailable, APX walks a configurable ordered list of fallback providers. Set the order and enable the router:
apx model enableapx model order ollama openrouter groqapx model status # probe all providers; show which would be used nowSee Configuration for the full super_agent.model_fallback schema.
Content-based routing
Section titled “Content-based routing”On top of the static fallback chain, opt-in per-turn routing rules (inspired by OpenHands’
RouterLLM) inspect the actual message — images, prompt/context size, channel, keywords — and
prefer a different model for that turn. Rules are evaluated in order; the first full match wins.
The preferred model is still health-checked and falls back down the regular chain if unavailable,
and an explicit per-request model override always beats the rules.
"super_agent": { "routing": { "enabled": true, "rules": [ { "model": "anthropic:claude-sonnet-5", "when": { "has_image": true } }, { "model": "groq:llama-3.3-70b-versatile", "when": { "max_prompt_chars": 200, "channels": ["telegram"] } }, { "model": "anthropic:claude-opus-4-8", "when": { "keywords": ["refactor", "architecture"] } } ] }}when conditions (all must hold): has_image, min_prompt_chars / max_prompt_chars,
min_context_chars, channels (list), keywords (case-insensitive substring list). An empty
when matches every turn — useful as a final catch-all.
- Memory system — how the agent retrieves relevant context across channels.
- Routines — schedule the super-agent or a plain
exec_agenton a cron. - Configuration — full
super_agent.*andidentity.jsonreference.