Skip to content

Super-agent

The super-agent is a mode, not a persona name. When you run apx exec without targeting a specific project agent, APX activates this mode: the daemon-level, tool-using loop that can reach every registered tool — file operations, MCP calls, Telegram, shell, tasks, sessions, and more.

The name the assistant uses when talking to you comes from ~/.apx/identity.json (default “APX”). References to “super-agent” are an internal technical term for the mode. You configure it with super_agent.* keys in ~/.apx/config.json, and routine kinds may still say super_agent — but users never see that label.

The mode is disabled by default. Set both keys before you first call it:

Terminal window
apx config set --global super_agent.enabled true
apx config set --global super_agent.model "anthropic:claude-sonnet-4-5"

Once enabled, every apx exec call without -a <agent> goes through this mode.

Terminal window
apx exec "<prompt>" # super-agent mode (default)
apx exec -- "<prompt>" # same; -- avoids flag parsing issues
apx exec -a reviewer "<prompt>" # named project agent instead
apx exec "<prompt>" --model gpt-5 # override the configured model for this call
apx exec "<prompt>" --max-tokens 2000 # cap output tokens
apx exec "<prompt>" --project my-app # pin to a specific project context
apx
$ apx exec 'list my open tasks across all projects'
Roby (super-agent) · anthropic:claude-sonnet-4-5

→ tool  tasks.list { project: "First Project", state: "open" }
→ tool  tasks.list { project: "Acme Store", state: "open" }
→ tool  tasks.list { project: "Data Pipeline", state: "open" }

You have 4 open tasks:
• First Project  — Wire Stripe webhooks before merge (due 06-16)
• Acme Store     — Retry-guard the flaky cart-total test (due 06-18)
• Acme Store     — Document the pairing flow for the web panel
• Data Pipeline  — Backfill ingestion CLI flag (due 06-20)

# 612 tok in / 198 out · 2.4s
apx exec in action — super-agent mode with tool calls visible

When the super-agent mode is active the daemon resolves the assistant’s name from ~/.apx/identity.json and displays it in status lines, TUI headers, and Telegram replies — never the string “super-agent”.

The configured identity is also a first-class A2A address. An agent can contact Roby as naturally as any project agent:

Terminal window
apx send crypto-analyst roby "BTC alert: decide whether the owner needs this on Telegram" --deliver

APX finds the sender’s project, files every identity alias under the stable super_agent peer, and runs the real super-agent loop: identity, relevant memory, project context, permission policy, and tools. This is different from apx exec, which represents an owner-authored turn. If the sender slug exists in several projects, add --project <name>.

Real project agents still win over aliases. If a project intentionally defines an agent whose slug is apx or whose name matches the configured identity, address selection chooses that agent inside its project.

Under the hood runSuperAgent() builds a system prompt (identity block + project inventory + memory block + active-threads block), selects a tool schema subset appropriate for the channel, then calls the engine via the standard runAgent() loop. The loop continues until the model stops calling tools or a signal aborts it.

The full registry has around 30 native tools plus registry-bridged tools (browser, fetch, search, glob, grep). The daemon picks a subset based on the channel to stay within cheap-tier TPM budgets:

ChannelTool setRationale
api / apx execFullDeliberate, user-picked model
routineFullScheduled, autonomous, user-picked model
web / codeFullLong-form workspace
telegramCore (~700 tokens)Cheap tiers, snappy replies
desktop / deckCoreSame

The model can always expand its surface by calling load_skill.

ToolWhat it does
create_task / list_tasks / get_taskAdd, list and read per-project TODO items
update_task / complete_task / comment_taskEdit one, move or close it, annotate its thread
call_agentHand work to a named project agent, in-process. It runs with its own tools, so it can do the work rather than describe it, and the exchange is filed as an agent-to-agent conversation you can read and continue
send_to_agentMessage another agent — a project agent, the super-agent, or a coding runtime — and get their answer. With background: true the work is left running and the sender is woken when it lands, so a long exchange no longer holds a turn. See Work left running
call_mcpCall any registered MCP server tool
call_runtimeDelegate to an external runtime (Claude Code / Codex …), synchronously or in the background
run_subagentSpawn an isolated sub-agent (fresh context, same tools minus user-interaction ones) that works a self-contained task to completion and returns its result — one nesting level max
run_shellExecute a shell command (gated by permission mode)
send_telegramSend a message via the Telegram plugin
send_fileHand you a file from disk — a screenshot, a generated image. See Sending you files
rememberWrite a durable fact to ~/.apx/memory.md
search_sessionsRetrieve past session transcripts
set_identityUpdate ~/.apx/identity.json fields

The super-agent can hand you a file with send_file: a browser screenshot, an image it generated, a render. The file arrives with the reply: in the web chat it shows inline with a download button, and on Telegram it goes out as a photo (or as a document for other kinds). Taking a screenshot does not send it by itself; the agent has to pass the saved path to send_file.

Only the kinds you enable are sent. Images are on by default, and everything else is off until you turn it on in ~/.apx/config.json:

{
"file_delivery": {
"kinds": { "image": true, "video": false, "audio": false, "document": false },
"max_mb": 20
}
}

document covers PDF, text, Markdown, CSV, JSON, Office files and .zip. Some things are never sent, whatever the config says: executables and scripts (.exe, .msi, .bat, .sh, .apk, .dmg, …), and anything under ~/.ssh, ~/.aws, ~/.gnupg or ~/.apx (except ~/.apx/media, where sent files are archived). An image whose bytes don’t match its extension is refused too, so renaming a file doesn’t get around the switch.

A peer’s turn is a full tool loop. It can take minutes, and send_to_agent waits for it by default — which is right only when the answer is the next thing the sender needs.

background: true changes that. The tool returns immediately with a job_id, the peer works on its own, and when it answers the sender is woken: the reply arrives as a new message on the same agent-to-agent thread, starting a fresh turn with the result in hand. Pass wake_me: false for a pure notification nobody needs an answer to.

The wake-up is an ordinary a2a message in the other direction, so the request and its answer live in one thread — there is nowhere new to look.

What the woken sender writes is filed as its own note, not as a reply: the peer already answered and is waiting on nothing. Filing it as a reply used to put “received, thanks” in the peer’s inbox, the peer confirmed back, and one answer turned into a round of acknowledgements, each a full tool loop.

  • Asking twice does not start it twice. While a hand-off to an agent is still running, calling that agent again (call_agent or send_to_agent) returns the running job — what it is doing and for how long — instead of opening a new turn. check_jobs reads the same status without touching anyone. A deliberate correction goes through with followup: true.
  • Nobody waits from inside an a2a thread. An agent answering another agent hands work on in the background, never by waiting on a third one; heavy work (a render, a production) runs as its own background job, not inside the conversation.
  • No question cards in a2a. Nobody there can answer one: if an agent needs the owner, the super-agent asks on its own channel.
  • The chat shows each hand-off as Waiting / Waited or Background, and the composer bar of the super-agent’s chat lists every agent working for it, with the chain below each one.

What the sender keeps, and what it does not

Section titled “What the sender keeps, and what it does not”

A woken turn is a new turn. It gets the reply, a recap of what was asked, and the thread’s own history — not the working context the sender had when it handed the job off. Anything needed to act on the answer belongs in the message itself.

A job that produced no answer wakes the sender too, and says so plainly: the peer failed, the job ran past its budget, or the daemon restarted while it ran and the work is unrecoverable. The sender is told not to report such a job as done.

LimitValueWhy
Jobs open per agent3Bounds how WIDE the fan-out goes
Hand-off depth3Bounds how DEEP it goes; counted across turns, wake-ups, call_agent and apx send
Tool steps per peer turn20 (super_agent.a2a_max_iters)A peer that runs out answers with what it did and what is left; the sender decides whether to continue
Job lifetime1 hour by defaultPast it the job is cut off and the sender told

Both limits come back as a message the agent can act on, never as a crash.

A turn nobody is watching — an a2a reply, a routine — also stops on a spent plan. When a provider answers that the account’s usage limit is reached (not a passing rate limit), that model cools down for an hour (super_agent.quota_cooldown_min), the unwatched turn ends instead of spending the rest of the fallback chain, and the owner gets one Telegram line saying so. A chat with a person in it still falls through to the next model. So does a model you chose — pinned to the agent, set on the routine, or the super-agent’s own model — when its fallback switch is on: that switch is your instruction to carry on down the router. The stop is for the router’s own account running dry under agents that inherit it.

Terminal window
curl -s localhost:7430/api/jobs?open=1 -H "Authorization: Bearer $TOKEN"

GET /api/jobs lists every job — ?open=1 for what is running now, plus ?project_id, ?from and ?status. GET /api/jobs/:id returns one. The panel reads the same route, and the live events feed carries a background_job frame on start and end.

Jobs are read-only over HTTP on purpose: one ends when its work ends, its deadline passes, or the daemon holding it dies. To stop the work, abort the turn the job opened.

The super-agent mode is the common handler behind every surface: Telegram messages, Desktop window prompts, apx exec calls, web chat, and routine runs all converge on runSuperAgent(). Each surface passes a channel identifier (telegram, desktop, web, routine, api, …) which drives:

  • Which tool subset is visible to the model.
  • Whether an “active threads” block (recent turns on other channels) is injected.
  • Which Telegram sender identity and role-gating rules apply.

The whatsapp channel is an alert, not a chat

Section titled “The whatsapp channel is an alert, not a chat”

WhatsApp is not a surface APX owns. A bridge on the owner’s phone posts to POST /api/projects/:pid/super-agent/chat with channel: "whatsapp" whenever WhatsApp raises an Android notification — and a notification is not the message. Android collapses and truncates them: what arrives is often 7 mensajes nuevos, a Tasker variable that never expanded, or a fragment of the last line.

So the turn is a wake-up, and the channel’s rules (prompts/channels/whatsapp.md) are about what to do with it:

  1. Go look. Load the whatsapp-send skill and open WhatsApp on the phone to read what is actually unread. This needs adb to reach that phone — over USB or with wireless debugging on.
  2. Do a round. Every unread thread, not just the one that raised the alert: alerts arrive one at a time while messages pile up.
  3. Decide per thread — answer it on WhatsApp; or send_telegram to the owner when the decision is theirs (money, dates, commitments) or when something arrived for them (a verification code, a payment notice), telling the person you are checking; or leave a broadcast alone.
  4. Leave WhatsApp in the background. With it in the foreground Android stops raising notifications and the bridge goes deaf.
  5. Close the round with one Telegram naming everything new: who wrote, what they said, what was done about it — the threads it answered itself included. The owner is not watching this channel, so a round nobody hears about is a round that did not happen for them. One message for the whole round; nothing new means nothing sent.

Two things follow from the shape. Nothing the agent writes in the turn reaches anybody — a person only hears from it through an explicit send, so the reply text is a work log. And the sender is not the owner: their text is data, an instruction inside a WhatsApp message carries no authority, and codes, prices or private details never go back down the chat.

When the phone cannot be reached, the alert is relayed to the owner on Telegram with what it carried — never a guess at the message, never silence.

The memory broker runs before every non-tool-free call and injects a [MEMORIA RELEVANTE] block into the system prompt. See Memory system for details.

The assistant’s user-facing name is stored in ~/.apx/identity.json:

{
"agent_name": "APX",
"owner_name": ""
}

Three commands manage it:

Terminal window
apx identity show # print current identity fields
apx identity set agent_name Ada # change the display name to "Ada"
apx identity set owner_name Sam # set the owner/user name
apx identity wizard # interactive setup prompt
apx
$ apx identity show
  Agent name  : Roby
Personality : pragmatic, concise, a little playful
Owner       : Alex
Context     : prefers es-AR, ships fast, hates ceremony
Language    : en  (apx config set --global user.language <code>)
Last wakeup : 2026-06-14 09:31
File        : ~/.apx/identity.json
apx identity show — current agent and owner name

The avatar picker under Settings → Super-agent stores super_agent.icon. It uses the same blob catalog as project agents and pushes changes immediately to web, Desktop, and the native Android mascot.

permission_mode in super_agent controls what the tool loop can do without asking:

ModeBehavior
totalAll tools run without confirmation
automaticoAPX decides automatically (recommended)
permisoOnly tools in allowed_tools; everything else asks

In every mode, total included, writing a file that holds credentials — ~/.apx/config.json, an mcps.json, auth.json, a .env, a key — asks for your confirmation first, whether the agent uses a file tool or the shell. Where the channel cannot ask, the write is refused. The native tools that manage those files (add_mcp, apx config set) are not affected.

Terminal window
apx permission set automatico

An opt-in layer on top of the permission mode (inspired by OpenHands’ LLMSecurityAnalyzer). When enabled, every tool schema gains a required security_risk field (LOW / MEDIUM / HIGH) that the model must fill as part of the call itself — no extra LLM pass. Calls graded at or above confirm_at pause for your confirmation on the active surface (web dialog, Telegram); a declined call comes back to the model as an error observation so it can re-plan.

"super_agent": {
"security_risk": {
"enabled": true,
"confirm_at": "HIGH",
"confirm_unknown": true
}
}
  • confirm_at — minimum grade that pauses (LOW confirms everything, HIGH only destructive/outward actions).
  • confirm_unknown — also pause when the model didn’t grade the call (weak models sometimes omit the field).

It’s not a duplicate of the permission mode — it gates by a different axis. Permission mode decides by tool identity (is this tool dangerous / on the allowlist); the risk analyzer decides by the model’s judgment of this specific action’s severity. So it adds a guard the permission mode can’t:

  • In automatico the risk grade replaces the static dangerous-tool confirmation.
  • In permiso the allowlist still applies and the risk gate composes on top.
  • In total it becomes a safety floor: everything runs free except a HIGH-graded action, which still stops for confirmation even under full trust. That’s the point — a catastrophic action gets caught no matter how much you trust the agent.

Verification loop (OpenHands critic pattern): an LLM judge scores the probability that the original request is fully satisfied. Below success_threshold the agent receives an internal verification note (what looks unmet, per the judge) and continues — up to max_iterations rounds. An unusable judge (engine down, unparseable reply) accepts the result rather than blocking it.

It covers two different ways a turn ends early:

  • enabled (opt-in) — a completion-contract turn (coding surfaces) declared itself done, and the judge checks that claim.
  • continue_unfinished (on by default) — a conversational turn (Telegram, web chat) stopped calling tools. That is also what a model does when it merely announces its next step (“ahora genero el SRT en inglés”) and writes no call: the task stops halfway and waits to be poked. The judge is the poke. A turn that ran no tools is never judged (chat has nothing to finish), nor is one that ended by asking you something — that one is waiting, not unfinished. “Asking you something” covers both ways a turn can ask: the ask_questions tool, and the reserved closing step that runs out of tool budget and offers to keep going. Continuing that one answers over your head, and since the next round hits the same wall it just repeats the same recap and the same question.

Project agents go through it too, not only the super-agent — through the continue_unfinished half only, since they never take the completion-contract path. It applies on every channel a project agent answers on, not just the watched ones: announcing a next step instead of taking it is the model’s failure, not the surface’s. The two exclusions above are what keep the cost down — after them the only turns judged are the ones that ran real tools and then went quiet without asking anything, which is one extra model call for a turn that would otherwise have sat there until you typed “seguí”.

"super_agent": {
"judge": {
"enabled": true,
"continue_unfinished": true,
"success_threshold": 0.6,
"max_iterations": 2,
"model": ""
}
}

model empty → the judge runs on super_agent.model, and failing that on the model the turn itself ran on — so a project agent is still judged on an install that never configured a super-agent. Set a cheap model to keep verification overhead low. Verdicts surface as judge_verdict events, on the result as result.judge, and on the Telegram message that closed the turn — so a turn that ran long can be traced back to the verdict that kept it going.

max_iterations caps how many times the judge may send a turn back — it does not multiply what the turn may spend. The rounds share the surface’s tool budget with the run that came before them: each gets what is left, and when too little is left to be worth a round (one action step plus the reserved closing one) the loop stops there and emits judge_budget_exhausted rather than starting it. So one message on the web chat costs at most web_max_iters, not (1 + max_iterations) × web_max_iters.

How many tool steps one turn may take depends on whether you can watch it work — and it is keyed on the surface, not on who is answering. A project agent in the web chat gets the same budget Roby does; a watched chat is a watched chat. (All four keys live under super_agent because that is where every tool-loop budget lives, including the routine one that governs project agents.)

SurfaceBudgetWhy
Telegramtelegram_max_iters (1000)Each real tool start is visible and Telegram keeps the typing indicator alive, so the turn normally runs until work is done.
Web chat + sidebarweb_max_iters (1000)Every call renders live and you’re one click from stopping it — you are the guardrail, so the turn runs until the work is done.
apx execcli_max_iters (40)Room to hand work to a project’s agent and check it; bounded, because scripts and other agents call it too. From apx exec a hand-off to another agent waits for the answer instead of running in the background.
Routines (non-Telegram)routine_max_iters (1000)Nobody is waiting mid-run to answer “want me to keep going?”.
Group roomsgroup_max_iters (50)The one surface where a single message fans out — see below.
Coding surfacesExplicit, plus the completion contractThey stop on finish, not on a count.

Each number is the budget for a turn, not for one pass of the tool loop: if the completion judge sends the turn back to finish something, those rounds spend what is left of the same number rather than being handed it again.

The three 1000 ceilings are runaway backstops, not normal stopping points: the loop already ends the moment the model stops calling tools, and stuck detection aborts the repeat-loops. Set any of these to 0 for the built-in default.

A group room is watched and stoppable like the web chat, so by the surface rule it would take the same ceiling — but one line you write can cascade into up to ten agent replies, each a full tool loop. Giving every speaker the 1:1 ceiling would quietly make one message worth ten of them. So a speaker gets a slice instead: room for genuine multi-step work, while the ten of them together still cost no more than one turn in your 1:1 chat with any of them. Same principle as the judge rounds above — the budget belongs to your turn, and everything it fans out into fits inside it. A speaker that runs out closes with what it got done and the room carries on; solo work that needs hundreds of steps belongs in that agent’s own chat, which has the ceiling for it.

"super_agent": {
"telegram_max_iters": 0,
"web_max_iters": 0,
"routine_max_iters": 0,
"group_max_iters": 0
}

When a budget is exhausted, the last step is reserved for a tool-free, model-authored closing message: what it did, what’s left, and whether to keep going. With the run-to-completion ceilings, that is a rare safety floor rather than the way normal work ends.

The tool loop watches for two loop shapes: the same call returning the same result action_repeat times, and the same call erroring error_repeat times in a row. On the first detection the model gets an in-band nudge (“you appear stuck — change approach or explain the blocker”); if it keeps looping, the turn is closed early with a model-authored wrap-up instead of burning the rest of the iteration budget. On by default:

"super_agent": {
"stuck_detection": { "enabled": true, "action_repeat": 4, "error_repeat": 3 }
}

If the primary model is unavailable, APX walks a configurable ordered list of fallback providers. Set the order and enable the router:

Terminal window
apx model enable
apx model order ollama openrouter groq
apx model status # probe all providers; show which would be used now

See Configuration for the full super_agent.model_fallback schema.

On top of the static fallback chain, opt-in per-turn routing rules (inspired by OpenHands’ RouterLLM) inspect the actual message — images, prompt/context size, channel, keywords — and prefer a different model for that turn. Rules are evaluated in order; the first full match wins. The preferred model is still health-checked and falls back down the regular chain if unavailable, and an explicit per-request model override always beats the rules.

"super_agent": {
"routing": {
"enabled": true,
"rules": [
{ "model": "anthropic:claude-sonnet-5", "when": { "has_image": true } },
{ "model": "groq:llama-3.3-70b-versatile", "when": { "max_prompt_chars": 200, "channels": ["telegram"] } },
{ "model": "anthropic:claude-opus-4-8", "when": { "keywords": ["refactor", "architecture"] } }
]
}
}

when conditions (all must hold): has_image, min_prompt_chars / max_prompt_chars, min_context_chars, channels (list), keywords (case-insensitive substring list). An empty when matches every turn — useful as a final catch-all.

  • Memory system — how the agent retrieves relevant context across channels.
  • Routines — schedule the super-agent or a plain exec_agent on a cron.
  • Configuration — full super_agent.* and identity.json reference.