A11 design¶
A research agent can stream its plan, investigate several briefs concurrently, keep intermediate evidence on the backend, and synthesize one report for its caller. A browser agent can call tools in the page that owns an editor or canvas. A support agent can retain structured conversation turns and resume them with another model provider.
The same runtime needs appear in responsive APIs: a live-captioning service sends transcript fragments while audio is arriving, an image-generation API reports progress before returning a PNG, and a model gateway serves tokens from a GPU process to browser, Python, and native clients.
These systems wait on independent producers, exchange structured and multimodal data, and move components between processes. An agent adds a model/tool loop; the runtime does not require one.
A11 keeps the operation's contract stable through those changes. The application describes work as actions and moves data through streams; storage and network placement remain deployment choices.
Research agent composition¶
The deep-research example presents a typical agent as
one deep-research action. Its inputs and outputs are the stable product
boundary: a topic comes in, while the plan, activity log, and report stream back
to a browser.
Inside, the Python handler calls a planning action and starts an investigation
for each brief. An asyncio.Semaphore bounds concurrency, and the included LLM
action handles every model call. The handler collects investigation reports
locally, then streams the synthesis through the public report node. The
intermediate reports do not cross the browser connection.
The composition uses the same mechanisms as a smaller tool-calling agent:
registered actions define capabilities, Interaction values carry model and
tool state, nodes move partial results, and a session reaches remote or
browser-hosted actions. The handler registers as one action, so callers only
need its topic, plan, and report ports. An optional Flow implementation can
express the same composition when it needs to be supplied at runtime.
Interface relationships¶
Solid arrows show what an interface owns or creates; dashed arrows show a dependency. The highlighted extension points have supported implementations in A11 and can also be supplied by an application. Hover over or focus a box for its role, then select it to open the corresponding reference.
Build streaming APIs and throughput-sensitive pipelines¶
An action is a useful API boundary whenever inputs or results have independent
lifecycles. A hosted model can accept configuration and prompt content on
separate inputs, then expose text, reasoning, usage, and the completed response
on separate outputs. A speech service can stream audio into recognition and
start translation on the first transcript fragment. A media endpoint can keep
progress as small JSON records while its final image remains encoded bytes with
an image/png media type.
These contracts remain modular because a caller routes each port to the reader that needs it and explicitly drains or omits the rest. Adding progress or diagnostics does not turn every result into a larger event union. The schema also remains the same for an in-process call, a model tool, and a service reached through a session.
A11 supports throughput by controlling data movement:
- streaming stages can begin before an upstream producer completes;
- backpressure propagates through bounded writers and stores;
- Flow places an explicit bound on concurrent work with
parallel N; - local node maps keep high-volume intermediates off the network;
- media types and serialization tags describe bytes without placing type names inside the payload;
- the native implementation supplies scheduling, storage, framing, and transport for Python and C++ services, while TypeScript shares the same wire and action contracts.
Deployment can keep decoding beside its caller, place inference on a GPU host, persist selected streams in SQLite or Redis, and move only the action ports that cross those boundaries. The generative-media, HTTP actions, and local-to-remote guides show these choices in complete APIs.
Connect models without rebuilding the application boundary¶
The included interact_with_llm action routes an Interaction to Claude,
Gemini, or Ollama. Text and reasoning stream on named output ports, while
completed interactions carry usage and conversation state. Application code
does not have to adopt each provider's event types or conversation envelope.
An Interaction retains structured text and images, tool calls, tool results,
token usage, and provider-native values needed to continue a turn. Actions
become model tools through their existing schemas, and the shared tool runner
turns requested calls back into ordinary A11 action calls. Switching an
included backend therefore does not require a second tool registry or
conversation data model.
Use a backend-specific action when the provider's controls are part of the
product contract, or interact_with_llm when the application chooses the
backend at runtime. Talk to a model starts with a streamed
turn; the LLM SDK articles cover portable state,
actions as tools, MCP, and tool execution.
Start with calls and streams, not an action graph¶
A11 does not require a graph object, scheduler, or checkpoint format to connect operations. Starting an action creates its input and output streams immediately. The action may produce output as soon as its inputs arrive, and a reader may start before the producer finishes. Waiting on data supplies the usual dependency ordering.
Arguments and results behave like asynchronous function parameters represented as named streams. They support linear pipelines, fan-out, loops, and parallel tool calls without a graph runtime.
Flow uses the same rule. Statements start concurrently unless an explicit ordering constraint says otherwise. A graph-shaped declaration is still possible: create the calls first, then connect their ports.
Flow source can arrive at runtime from an application, user, or model. The host resolves it against its current action registry, reports syntax and contract errors before dispatch, and restricts model-authored calls with the same per-turn allow-list as direct tool calls. A flow cannot import code or discover capabilities outside the registry. This provides dynamic composition within an explicit set of operations and types.
search = run web-search(query: question, limit: 3)
brief = run llm-summarize(question: question)
for hit in search.hits parallel 3 {
page = run web-fetch(url: hit.url)
page.text | truncate 2000 -> brief.pages
}
brief.summary -> answer
brief starts with its pages port open. Fetches feed that port as results
arrive, and the runtime closes it when the loop finishes. Data arrival
coordinates the search, fetch, and summary steps.
Intermediate results travel directly between ports. They need not be copied by a model, added to its context, collected into a graph state object, or sent to the dispatching peer. Runtime composition therefore retains the streaming and placement choices of the underlying actions.
Build a capability once¶
An Action is a named operation with described
input and output ports. Its handler can run in the caller's process or be
registered on another peer. Callers feed and read the same ports in either
case.
The deployment selects the execution location:
- keep a formatter or parser beside its caller;
- put model access behind a service that owns credentials or a GPU;
- run a browser action in the page that owns the canvas or editor state;
- compose several registered actions into one workflow.
The same ActionSchema also supplies the
function definition shown to a model and the contract advertised to a remote
peer. Tool adapters and MCP integrations translate at this boundary, so the
application does not maintain separate schemas for local calls, model tool
calling, and service dispatch.
The local-to-remote guide follows one action across that boundary. Browser-hosted tools show the reverse direction: a backend dispatches an action to the connected page.
Discover operations at runtime and reuse connections¶
gRPC commonly starts with a service definition and generated client and server code. Its unary and streaming RPCs run through reusable channels. This model, descended from Google's Stubby, is effective when services have a stable interface and every client can build against it.
A11 retains efficient connection reuse but makes the operation layer dynamic.
A registry exposes its current ActionSchema values at runtime, and callers
bind named node streams without generating a client stub. This fits model
tools, plugins, browser capabilities, and tenant-specific services whose
available operations can change while the application is running.
The dynamic contract does not require a physical connection for every call or
port. Attaching a WireStream starts or accepts that transport immediately.
The Session then multiplexes action control, input fragments, output
fragments, and shutdown messages over each active stream. Many logical calls
and nodes can therefore share a small number of network connections, with
explicit limits, backpressure, deadlines, cancellation, and status handling.
See the Session lifecycle for the exact routing and
completion rules.
Return work as it is produced¶
Every action port is an AsyncNode, an ordered
stream of values. A handler can write progress, tokens, audio frames, records,
or one complete object. The caller chooses whether to process each value or
collect a unary result.
put()appends a value and exposes confirmation from the backing store.async foriterates over values until the stream ends.consume()reads a port expected to contain one complete value.finalize()marks the logical end of the data and closes the writer.abort_with_status()ends the stream with a structured failure.
Because a reader can start before the producer finishes, connected actions can overlap their work. A text interface can display model output immediately; a media action can report progress on one port while preparing an image on another. See streaming through a node and separate progress and result ports.
Choose where stream data lives¶
A ChunkStore is the extensible storage
interface behind a node's ordered chunks. A11 includes an in-memory store,
embedded SQLite persistence, and Redis-backed streams. Applications can provide
a ChunkStore and factory for another database, object store, retention policy,
or test environment without changing the producer, reader, or action schema:
- SQLite can retain conversations and reopen them after a process restart;
- Redis can connect programs through durable, named streams even when their lifetimes do not overlap;
- a custom store can apply an application's retention, persistence, or testing policy.
A11 implements node ordering, cursors, and final status over established storage systems. SQLite supplies embedded transactions and local durability; Redis supplies shared streams, atomic operations, deployment tooling, and its own replication and scaling options. Applications can use infrastructure they already operate instead of adopting a dedicated A11 storage service. The backend's operational limits still apply: SQLite is primarily a one-machine choice, while Redis Cluster needs the routing setup described in the Redis reference.
The persistent chat and Redis stream guides demonstrate these choices.
This covers many cases for which an agent framework would introduce a separate
conversation-memory or stream-checkpoint layer. Store the completed
Interaction values when conversation history is the requirement; choose a
durable node store when live stream data must survive process boundaries. These
are explicit data-retention choices, not hidden state in an agent runner.
Connect peers when they need live calls¶
A Session dispatches actions and carries their
node data over one or more connections. It tracks in-flight work so callers and
services can drain and close cleanly.
The connection implements the extensible
WireStream interface. A11 includes
in-process, WebSocket, HTTP Server-Sent Events, and WebRTC implementations.
Applications can implement the same interface for another bidirectional
transport while sessions, actions, and nodes remain unchanged. The
echo service shows the complete client and server
lifecycle; the browser client uses an
HTTP-compatible transport.
Use a session for live action dispatch between peers. Use a shared store when the main requirement is durable stream data that either side may read later.
Make completion and failure observable¶
Stream termination is part of the data contract. Finalization tells a reader that it received a complete result; an aborted stream carries a status instead of appearing to be valid but truncated.
Actions, sessions, and connections also expose completion explicitly. A caller can await the result it needs, wait for the full action, or let a context manager drain a connection during shutdown. Deadlines and cancellation propagate through nested action calls.
The lifecycle articles describe the exact transitions for nodes, actions, sessions, and connections.
Use the language suited to each boundary¶
Python, TypeScript, and C++ expose the same action, node, session, storage, and transport concepts. A Python service, TypeScript browser, and native component can share action schemas and exchange the same wire messages while each uses the conventions of its language.
Start with the examples by task, or go directly to the Python, TypeScript, or C++ reference.