Skip to content

Run a model in the browser

The Browser clients guide calls a model hosted on a server. Run the model locally with interact_with_gemma, which loads a Gemma-family model into the page and runs it on the browser's GPU through WebGPU. Prompts and responses stay in the page, and the reply streams through an AsyncNode.

interact_with_gemma is an A11 action with the same interaction ports as the other LLM backends: an interactions input, a unary config input, and text_output / new_interactions outputs. The caller uses the same port contract as interact_with_llm; its handler executes through WebGPU.

Before you start

Install a11@npm:@curiositystack/a11. Use a browser with WebGPU enabled and provide a hosted Gemma model asset (.task or .litertlm) that MediaPipe can load. Serve the large model file with permissive CORS. This setup does not require an API key or model server.

Try it

Paste a hosted Gemma model URL and send a message. The first message downloads and compiles the model in the browser. Later generation is local, and each reply streams token by token. The browser cache avoids another download after a reload. A WebGPU-capable browser is required.

1. Import the action contract

Import the backend and SDK types. INTERACT_WITH_GEMMA_SCHEMA describes the ports and registers like any other schema.

import {
    Action,
    ActionRegistry,
    INTERACT_WITH_GEMMA_SCHEMA,
    interactWithGemma,
    makeTextMessageInteraction,
    parseInteraction,
    isOk,
    StatusCode,
    type Status,
} from '@curiositystack/a11';

const need = <T>(value: T | Status): T => {
    if (!isOk(value)) throw new Error(`${StatusCode[value.code]}: ${value.message}`);
    return value as T;
};

2. Register and run the action locally

There is no session and no transport. Create the action from its schema, bind the handler, and run() it. run() starts the handler on the same node map, so the ports opened next are the ones the handler reads and writes.

const registry = new ActionRegistry();
need(registry.register('interact_with_gemma', INTERACT_WITH_GEMMA_SCHEMA, interactWithGemma));

const action = need(Action.create(INTERACT_WITH_GEMMA_SCHEMA, {
    handler: interactWithGemma,
    registry,
}));
need(action.run());

3. Feed the conversation and the model URL

The interactions port takes the whole conversation; the unary config port carries the browser-specific knobs, most importantly model_asset_path — the URL of the Gemma model to download and run. makeTextMessageInteraction builds a portable text turn.

const user = need(await makeTextMessageInteraction('Explain WebGPU in one sentence.'));

const interactions = need(await action.getInput('interactions'));
need(await interactions.finalize(user));

const config = need(await action.getInput('config'));
need(await config.finalize({}));

The default config loads Gemma 3n E2B from Hugging Face. Set model_asset_path to select another compatible asset. The asset is fetched with redirects followed (the resolve URL 302s to a CDN), then handed to the runtime as bytes. The downloaded bytes are stored in the browser's Cache Storage, so a page reload serves the model from disk instead of downloading it again. The first turn triggers the download and WebGPU compilation; later turns reuse the loaded model.

4. Stream the reply as it arrives

text_output is an AsyncNode. Reading it in a loop lets tokens appear the moment the model produces them — next() returns each streamed piece and null once the turn is complete.

const output = need(await action.getOutput('text_output', false));
let reply = '';
while (true) {
    const token = need(await output.next({timeoutMs: 120_000}));
    if (token === null) break;
    reply += token;
    render(reply);
}

The completed assistant turn also lands, structured, on new_interactions. Keep it and prepend it to the next interactions write to continue the conversation:

const newInteractions = need(await action.getOutput('new_interactions', false));
const assistant = need(parseInteraction(need(await newInteractions.next())));
need(await action.wait(5_000));
history = [...history, user, assistant];

5. Provide another runtime (optional)

By default, the handler dynamically imports Google's MediaPipe LlmInference task and runs it on WebGPU. Use setGemmaEngineFactory to cache a loaded model across turns, report download progress, or provide another engine:

import {setGemmaEngineFactory, type GemmaEngine} from '@curiositystack/a11';

setGemmaEngineFactory(async (config) => {
    const engine: GemmaEngine = {
        async generate(prompt, onToken) {
            return '...';
        },
    };
    return engine;
});

A factory returns a StatusOr<GemmaEngine>. It reports failures with a status such as unavailableError(...); the runtime aborts the output ports with that status.

6. Display failures

Every A11 call returns a StatusOr<T>: success values pass isOk, failures carry a code, message, and details. Convert them at the UI boundary — a missing model URL, an absent WebGPU adapter, or a load error all arrive as ordinary statuses, not exceptions:

try {
    await runTurn(text);
} catch (error) {
    errorRegion.textContent = error instanceof Error ? error.message : String(error);
}