Skip to content

Run actions requested by an interaction

When a model requests tools, its assistant Interaction contains action_calls and the corresponding action_inputs. The included runner turns those records back into live nested A11 actions, waits for them, and returns output fragments. Those outputs belong on a new user interaction, which follows the assistant interaction containing the calls.

This is the execution part of the agent loop. It handles parallel tool calls, deadlines, partial failure, and output collection without coupling an action handler to a provider SDK.

import a11

from a11.sdk.llm import Interaction, Role
from a11.sdk.llm_tools.runner import execute_actions_from_interaction

# `parent_action` is the running model action. Its registry and deadline are
# inherited by each nested tool call.
outputs = await execute_actions_from_interaction(
    assistant_interaction,
    parent_action,
)

tool_result_interaction = Interaction(
    previous_interaction_id=assistant_interaction.id,
    role=Role.USER,
    action_outputs=outputs,
    content=[
        a11.to_chunk(
            {
                "role": "user",
                # This adapter returns the provider's tool-result blocks.
                "content": await build_tool_result_blocks(outputs),
            }
        )
    ],
)

Do not mutate assistant_interaction.action_outputs. The assistant interaction is the model's request to use tools; the new user interaction is the application's response. Keeping them as separate, ordered turns preserves the conversation structure expected by Claude and the other included handlers.

The runner handles five lifecycle responsibilities for each requested action:

  1. It checks every requested name against x-a11-allowed-llm-actions.
  2. It resolves the schema and handler from the action registry.
  3. It creates a nested action, preserves the model’s call ID, and propagates the parent deadline.
  4. It streams input fragments to their ports while leaving autofilled ports to the runtime.
  5. It waits for completion and maps output port names into the JSON-facing fields declared by the action schema.

The runner returns those fragments; it does not create the user interaction or encode provider content. For example, interact_with_claude builds Anthropic tool_result blocks, places them in a new Role.USER interaction alongside action_outputs, emits that interaction, and only then asks Claude to continue.

Image outputs

An output port declaring an image/* media type carries the encoding itself, which has no JSON form. build_tool_results splits those fragments out and hands them to the backend as NormalizedContentType.IMAGE parts, so the result document holds the action's other ports and the frames travel as pictures.

from a11.sdk.llm import decoded_output_content

text, images = await decoded_output_content(outputs["call-1"])
text      # '{"legend": {"floor_index": 0, ...}}'
images    # [NormalizedPart(type=IMAGE, mime_type="image/png", data="iVBOR...")]

Where each provider puts them:

backend the frames ride on
interact_with_claude an image block inside the tool_result
interact_with_claude_code MCP ImageContent in the tool result
interact_with_codex a Codex --image attachment on the resumed turn
interact_with_gemini a user_input step after the function_result
interact_with_gpt a user message's image_url part, after the tool message
interact_with_ollama a user message's images, after the tool message
interact_with_vllm a user message's image_url part, after the tool message

format_result may answer with a list: the three routes whose tool message is text-only send two messages for one call.

A port carrying bytes under any other media type stays in the result document, base64url-encoded. A msgpack payload and a packed grid are data a model reads about.

Narrate a run without telling the model

A handler reports user-visible activity through log. A log can contain a summary followed by supporting detail. The shell tools in a11.sdk.bash use this structure.

await action.log("Ran `git status`.\n\n2 lines of output.")
await action.log("resolved the shell", internal=True)   # A11's own bookkeeping

The schema does not declare a port for this log, and callers do not drain it. The runner reads it separately from action outputs. Model tool results include only declared output ports, while user-visible logs omit entries marked internal.

executed = await execute_actions_from_interaction(assistant_interaction, action)

executed.logs            # {tool call id: log}, only for calls that narrated
executed.log_metadata()  # {"tool_logs": b'{"call id": "..."}'}, or {}

Merge log_metadata() into the backend_specific_metadata of the interaction carrying the tool results, which is what the included handlers do. Metadata is the one part of an interaction no backend turns into provider content, so the log stays out of the model's context while still being stored with the conversation and available to a replayed transcript.

Make the allow-list explicit

The registry defines what the application can run. The header defines what this particular model call may run. Both are required boundaries.

from a11.sdk.llm import LlmHeaders

parent_action.set_header(
    LlmHeaders.ALLOWED_LLM_ACTIONS.value,
    # Patterns are full-match regular expressions.
    r"look_up_order|search_catalog",
)

Do not build this value from untrusted model output. Choose it from application policy, user permissions, and the needs of the current workflow. A call outside the allow-list fails with PERMISSION_DENIED before an action is started.

Most applications do not call execute_actions_from_interaction directly: the included interact_with_* handlers call it while continuing the model/tool loop. Use it directly when building a custom model provider or an interaction orchestrator with its own loop.