Skip to content

A11 design

A research agent can stream its plan, investigate several briefs concurrently, keep intermediate evidence on the backend, and synthesize one report for its caller. A browser agent can call tools in the page that owns an editor or canvas. A support agent can retain structured conversation turns and resume them with another model provider.

The same runtime needs appear in responsive APIs: a live-captioning service sends transcript fragments while audio is arriving, an image-generation API reports progress before returning a PNG, and a model gateway serves tokens from a GPU process to browser, Python, and native clients.

These systems wait on independent producers, exchange structured and multimodal data, and move components between processes. An agent adds a model/tool loop; the runtime does not require one.

A11 keeps the operation's contract stable through those changes. The application describes work as actions and moves data through streams; storage and network placement remain deployment choices.

Research agent composition

The deep-research example presents a typical agent as one deep-research action. Its inputs and outputs are the stable product boundary: a topic comes in, while the plan, activity log, and report stream back to a browser.

Inside, the Python handler calls a planning action and starts an investigation for each brief. An asyncio.Semaphore bounds concurrency, and the included LLM action handles every model call. The handler collects investigation reports locally, then streams the synthesis through the public report node. The intermediate reports do not cross the browser connection.

The composition uses the same mechanisms as a smaller tool-calling agent: registered actions define capabilities, Interaction values carry model and tool state, nodes move partial results, and a session reaches remote or browser-hosted actions. The handler registers as one action, so callers only need its topic, plan, and report ports. An optional Flow implementation can express the same composition when it needs to be supplied at runtime.

Interface relationships

Solid arrows show what an interface owns or creates; dashed arrows show a dependency. The highlighted extension points have supported implementations in A11 and can also be supplied by an application. Hover over or focus a box for its role, then select it to open the corresponding reference.

A11 interfaces and their relationships A navigable map from Flow and model integrations through actions, sessions, nodes, storage, transport, and data types. Flow Flow checks and executes runtime compositions against the host's current ActionRegistry; every flow is also an Action. LLM SDK Provider adapters expose model requests, streamed outputs, and tool calls through Actions and AsyncNodes. Service Service shares one action registry across listeners and creates a Session for each connected peer. WireStream EXTENSIBLE Use the in-process, WebSocket, HTTP SSE, or WebRTC implementations, or implement WireStream for another bidirectional transport. ActionRegistry ActionRegistry is the runtime catalogue of schemas and handlers used for discovery, local calls, remote dispatch, tools, and Flow. Session Session resolves inbound actions, owns the peer's NodeMap, and routes many action and node streams over a bounded set of transports. Data & serialization Chunks carry bytes, media metadata, and status through stores and transports; the serialization registry maps application values to them. Action Action binds an ActionSchema to named input and output AsyncNodes, then runs a local handler or dispatches the call through a Session. NodeMap NodeMap gives actions and incoming fragments a shared node namespace inside a Session. AsyncNode AsyncNode is every action port: producers append values, readers iterate streams or read unary values; final status makes completion observable. ChunkStore EXTENSIBLE Use local memory, embedded SQLite, or shared Redis, or implement ChunkStore for application-specific retention and placement.

Build streaming APIs and throughput-sensitive pipelines

An action is a useful API boundary whenever inputs or results have independent lifecycles. A hosted model can accept configuration and prompt content on separate inputs, then expose text, reasoning, usage, and the completed response on separate outputs. A speech service can stream audio into recognition and start translation on the first transcript fragment. A media endpoint can keep progress as small JSON records while its final image remains encoded bytes with an image/png media type.

These contracts remain modular because a caller routes each port to the reader that needs it and explicitly drains or omits the rest. Adding progress or diagnostics does not turn every result into a larger event union. The schema also remains the same for an in-process call, a model tool, and a service reached through a session.

A11 supports throughput by controlling data movement:

  • streaming stages can begin before an upstream producer completes;
  • backpressure propagates through bounded writers and stores;
  • Flow places an explicit bound on concurrent work with parallel N;
  • local node maps keep high-volume intermediates off the network;
  • media types and serialization tags describe bytes without placing type names inside the payload;
  • the native implementation supplies scheduling, storage, framing, and transport for Python and C++ services, while TypeScript shares the same wire and action contracts.

Deployment can keep decoding beside its caller, place inference on a GPU host, persist selected streams in SQLite or Redis, and move only the action ports that cross those boundaries. The generative-media, HTTP actions, and local-to-remote guides show these choices in complete APIs.

Connect models without rebuilding the application boundary

The included interact_with_llm action routes an Interaction to Claude, Gemini, or Ollama. Text and reasoning stream on named output ports, while completed interactions carry usage and conversation state. Application code does not have to adopt each provider's event types or conversation envelope.

An Interaction retains structured text and images, tool calls, tool results, token usage, and provider-native values needed to continue a turn. Actions become model tools through their existing schemas, and the shared tool runner turns requested calls back into ordinary A11 action calls. Switching an included backend therefore does not require a second tool registry or conversation data model.

Use a backend-specific action when the provider's controls are part of the product contract, or interact_with_llm when the application chooses the backend at runtime. Talk to a model starts with a streamed turn; the LLM SDK articles cover portable state, actions as tools, MCP, and tool execution.

Start with calls and streams, not an action graph

A11 does not require a graph object, scheduler, or checkpoint format to connect operations. Starting an action creates its input and output streams immediately. The action may produce output as soon as its inputs arrive, and a reader may start before the producer finishes. Waiting on data supplies the usual dependency ordering.

Arguments and results behave like asynchronous function parameters represented as named streams. They support linear pipelines, fan-out, loops, and parallel tool calls without a graph runtime.

Flow uses the same rule. Statements start concurrently unless an explicit ordering constraint says otherwise. A graph-shaped declaration is still possible: create the calls first, then connect their ports.

Flow source can arrive at runtime from an application, user, or model. The host resolves it against its current action registry, reports syntax and contract errors before dispatch, and restricts model-authored calls with the same per-turn allow-list as direct tool calls. A flow cannot import code or discover capabilities outside the registry. This provides dynamic composition within an explicit set of operations and types.

search = run web-search(query: question, limit: 3)
brief  = run llm-summarize(question: question)

for hit in search.hits parallel 3 {
  page = run web-fetch(url: hit.url)
  page.text | truncate 2000 -> brief.pages
}

brief.summary -> answer

brief starts with its pages port open. Fetches feed that port as results arrive, and the runtime closes it when the loop finishes. Data arrival coordinates the search, fetch, and summary steps.

Intermediate results travel directly between ports. They need not be copied by a model, added to its context, collected into a graph state object, or sent to the dispatching peer. Runtime composition therefore retains the streaming and placement choices of the underlying actions.

Build a capability once

An Action is a named operation with described input and output ports. Its handler can run in the caller's process or be registered on another peer. Callers feed and read the same ports in either case.

The deployment selects the execution location:

  • keep a formatter or parser beside its caller;
  • put model access behind a service that owns credentials or a GPU;
  • run a browser action in the page that owns the canvas or editor state;
  • compose several registered actions into one workflow.

The same ActionSchema also supplies the function definition shown to a model and the contract advertised to a remote peer. Tool adapters and MCP integrations translate at this boundary, so the application does not maintain separate schemas for local calls, model tool calling, and service dispatch.

The local-to-remote guide follows one action across that boundary. Browser-hosted tools show the reverse direction: a backend dispatches an action to the connected page.

Discover operations at runtime and reuse connections

gRPC commonly starts with a service definition and generated client and server code. Its unary and streaming RPCs run through reusable channels. This model, descended from Google's Stubby, is effective when services have a stable interface and every client can build against it.

A11 retains efficient connection reuse but makes the operation layer dynamic. A registry exposes its current ActionSchema values at runtime, and callers bind named node streams without generating a client stub. This fits model tools, plugins, browser capabilities, and tenant-specific services whose available operations can change while the application is running.

The dynamic contract does not require a physical connection for every call or port. Attaching a WireStream starts or accepts that transport immediately. The Session then multiplexes action control, input fragments, output fragments, and shutdown messages over each active stream. Many logical calls and nodes can therefore share a small number of network connections, with explicit limits, backpressure, deadlines, cancellation, and status handling. See the Session lifecycle for the exact routing and completion rules.

Return work as it is produced

Every action port is an AsyncNode, an ordered stream of values. A handler can write progress, tokens, audio frames, records, or one complete object. The caller chooses whether to process each value or collect a unary result.

  • put() appends a value and exposes confirmation from the backing store.
  • async for iterates over values until the stream ends.
  • consume() reads a port expected to contain one complete value.
  • finalize() marks the logical end of the data and closes the writer.
  • abort_with_status() ends the stream with a structured failure.

Because a reader can start before the producer finishes, connected actions can overlap their work. A text interface can display model output immediately; a media action can report progress on one port while preparing an image on another. See streaming through a node and separate progress and result ports.

Choose where stream data lives

A ChunkStore is the extensible storage interface behind a node's ordered chunks. A11 includes an in-memory store, embedded SQLite persistence, and Redis-backed streams. Applications can provide a ChunkStore and factory for another database, object store, retention policy, or test environment without changing the producer, reader, or action schema:

  • SQLite can retain conversations and reopen them after a process restart;
  • Redis can connect programs through durable, named streams even when their lifetimes do not overlap;
  • a custom store can apply an application's retention, persistence, or testing policy.

A11 implements node ordering, cursors, and final status over established storage systems. SQLite supplies embedded transactions and local durability; Redis supplies shared streams, atomic operations, deployment tooling, and its own replication and scaling options. Applications can use infrastructure they already operate instead of adopting a dedicated A11 storage service. The backend's operational limits still apply: SQLite is primarily a one-machine choice, while Redis Cluster needs the routing setup described in the Redis reference.

The persistent chat and Redis stream guides demonstrate these choices.

This covers many cases for which an agent framework would introduce a separate conversation-memory or stream-checkpoint layer. Store the completed Interaction values when conversation history is the requirement; choose a durable node store when live stream data must survive process boundaries. These are explicit data-retention choices, not hidden state in an agent runner.

Connect peers when they need live calls

A Session dispatches actions and carries their node data over one or more connections. It tracks in-flight work so callers and services can drain and close cleanly.

The connection implements the extensible WireStream interface. A11 includes in-process, WebSocket, HTTP Server-Sent Events, and WebRTC implementations. Applications can implement the same interface for another bidirectional transport while sessions, actions, and nodes remain unchanged. The echo service shows the complete client and server lifecycle; the browser client uses an HTTP-compatible transport.

Use a session for live action dispatch between peers. Use a shared store when the main requirement is durable stream data that either side may read later.

Make completion and failure observable

Stream termination is part of the data contract. Finalization tells a reader that it received a complete result; an aborted stream carries a status instead of appearing to be valid but truncated.

Actions, sessions, and connections also expose completion explicitly. A caller can await the result it needs, wait for the full action, or let a context manager drain a connection during shutdown. Deadlines and cancellation propagate through nested action calls.

The lifecycle articles describe the exact transitions for nodes, actions, sessions, and connections.

Use the language suited to each boundary

Python, TypeScript, and C++ expose the same action, node, session, storage, and transport concepts. A Python service, TypeScript browser, and native component can share action schemas and exchange the same wire messages while each uses the conventions of its language.

Start with the examples by task, or go directly to the Python, TypeScript, or C++ reference.