Run a model in the browser¶
The Browser clients guide calls a model hosted on a
server. Run the model locally with interact_with_gemma, which loads a
Gemma-family model into the page and runs it
on the browser's GPU through
WebGPU. Prompts and
responses stay in the page, and the reply streams through an AsyncNode.
interact_with_gemma is an A11 action with the same interaction ports as the
other LLM backends: an interactions input, a unary config input, and
text_output / new_interactions outputs. The caller uses the same port
contract as interact_with_llm; its handler executes through WebGPU.
Before you start
Install a11@npm:@curiositystack/a11. Use a browser with WebGPU enabled and
provide a hosted Gemma model asset (.task or .litertlm) that MediaPipe
can load. Serve the large model file with permissive CORS. This setup does
not require an API key or model server.
Try it¶
Paste a hosted Gemma model URL and send a message. The first message downloads and compiles the model in the browser. Later generation is local, and each reply streams token by token. The browser cache avoids another download after a reload. A WebGPU-capable browser is required.
1. Import the action contract¶
Import the backend and SDK types. INTERACT_WITH_GEMMA_SCHEMA describes the
ports and registers like any other schema.
import {
Action,
ActionRegistry,
INTERACT_WITH_GEMMA_SCHEMA,
interactWithGemma,
makeTextMessageInteraction,
parseInteraction,
isOk,
StatusCode,
type Status,
} from '@curiositystack/a11';
const need = <T>(value: T | Status): T => {
if (!isOk(value)) throw new Error(`${StatusCode[value.code]}: ${value.message}`);
return value as T;
};
2. Register and run the action locally¶
There is no session and no transport. Create the action from its schema, bind
the handler, and run() it. run() starts the handler on the same node map, so
the ports opened next are the ones the handler reads and writes.
const registry = new ActionRegistry();
need(registry.register('interact_with_gemma', INTERACT_WITH_GEMMA_SCHEMA, interactWithGemma));
const action = need(Action.create(INTERACT_WITH_GEMMA_SCHEMA, {
handler: interactWithGemma,
registry,
}));
need(action.run());
3. Feed the conversation and the model URL¶
The interactions port takes the whole conversation; the unary config port
carries the browser-specific knobs, most importantly model_asset_path — the
URL of the Gemma model to download and run. makeTextMessageInteraction
builds a portable text turn.
const user = need(await makeTextMessageInteraction('Explain WebGPU in one sentence.'));
const interactions = need(await action.getInput('interactions'));
need(await interactions.finalize(user));
const config = need(await action.getInput('config'));
need(await config.finalize({}));
The default config loads Gemma 3n E2B from Hugging Face. Set
model_asset_path to select another compatible asset. The asset is fetched
with redirects followed (the
resolve URL 302s to a CDN), then handed to the runtime as bytes. The
downloaded bytes are stored in the browser's Cache
Storage, so a page
reload serves the model from disk instead of downloading it again. The first
turn triggers the download and WebGPU compilation; later turns reuse the loaded
model.
4. Stream the reply as it arrives¶
text_output is an AsyncNode. Reading it in a loop lets tokens appear the
moment the model produces them — next() returns each streamed piece and
null once the turn is complete.
const output = need(await action.getOutput('text_output', false));
let reply = '';
while (true) {
const token = need(await output.next({timeoutMs: 120_000}));
if (token === null) break;
reply += token;
render(reply);
}
The completed assistant turn also lands, structured, on new_interactions.
Keep it and prepend it to the next interactions write to continue the
conversation:
const newInteractions = need(await action.getOutput('new_interactions', false));
const assistant = need(parseInteraction(need(await newInteractions.next())));
need(await action.wait(5_000));
history = [...history, user, assistant];
5. Provide another runtime (optional)¶
By default, the handler dynamically imports Google's MediaPipe LlmInference
task and runs it on WebGPU. Use setGemmaEngineFactory to cache a loaded model
across turns, report download progress, or provide another engine:
import {setGemmaEngineFactory, type GemmaEngine} from '@curiositystack/a11';
setGemmaEngineFactory(async (config) => {
const engine: GemmaEngine = {
async generate(prompt, onToken) {
return '...';
},
};
return engine;
});
A factory returns a StatusOr<GemmaEngine>. It reports failures with a status
such as unavailableError(...); the runtime aborts the output ports with that
status.
6. Display failures¶
Every A11 call returns a StatusOr<T>: success values pass isOk, failures
carry a code, message, and details. Convert them at the UI boundary — a missing
model URL, an absent WebGPU adapter, or a load error all arrive as ordinary
statuses, not exceptions: