Speeding Up Generative UI with Async A2UI Caching
LLM startup latency kills generative UI. Learn how to pre-generate A2UI messages on the backend, cache them, and render instant GenUI in Flutter on app launch.
Published on • October 2, 2026
AI Assistant

Generative UI promises an interface that adapts to user intent in real time. But it introduces a familiar problem: startup latency. If your app calls a large language model on launch to build the first screen, users wait while the model processes system prompts, fetches records, and constructs UI payloads.
The fix is to decouple UI generation from the client’s runtime lifecycle: pre-generate A2UI messages in the background, cache them, and load them instantly at startup.
Source: Speeding up generative UI with async A2UI, The Flutter Blog, Aug 13, 2026.
The pattern
The reference implementation is Commis, a sample assistant for commercial kitchens. Its UX centers on a chat interface where a chef coordinates catering jobs — but when the chef opens the app, they shouldn’t start from a blank chat. The agent should offer pre-generated UI with info and actions relevant to each job.
The architecture separates two phases:
- Backend generation — whenever data changes, a function generates the UI and writes the A2UI JSON to a database collection.
- Client consumption — on launch, the Flutter app pulls the pre-computed UI straight from the database, bypassing the LLM entirely, and feeds it into the
genuitransport layer.
1. Asynchronous UI generation on the backend
The server half uses the experimental Dart support in Firebase Cloud Functions with a Firestore trigger. When a record changes, the function reads the data, prompts Gemini to construct the UI, and writes the result:
firebase.firestore.onDocumentWritten((event) async {
final jobId = event.params['jobId'];
final jobData = event.data?.after?.data();
if (jobData == null) return;
try {
final generator = ContentGenerator(apiKey);
final message = await generator.generateFeed(jobData.toString());
if (message == noUiSentinel) {
await firestore.collection('feed').doc(jobId).delete();
return;
}
await firestore.collection('feed').doc(jobId).set({'message': message});
} catch (e) {
print('Error generating UI: $e');
}
}, document: 'jobs/{jobId}');
Two details matter here:
- A restrictive system prompt. The sample instructs the model: “RESPOND ONLY WITH A2UI MESSAGES FOR A NavigationCard OR NO GENERATED UI. NOTHING ELSE.” You can strip the guardrails down or hand the agent your entire catalog — that’s a choice, not a requirement.
- A sentinel for “no UI.” If a job is far in the future, the model returns a sentinel string and the function clears the cached document. Caching without invalidation is just stale state.
In the sample, near-term jobs produce a map-and-navigation card; distant jobs produce nothing.
2. Consuming cached UI in Flutter
On the client, initialization fetches cached messages and routes them two ways:
Future<void> _initAgent() async {
final repository = context.read<FirestoreRepository>();
String? feeds;
try {
final jobs = await repository.getJobs();
final feedFutures = jobs.map((job) => repository.getFeedMessage(job.id));
final results = await Future.wait(feedFutures);
feeds = results
.map((r) => r?.trim() ?? '')
.where((r) => r.isNotEmpty)
.join('\n\n');
} catch (e) {
debugPrint('Error initializing agent: $e');
}
_agentService = FirebaseAILogicService(
repository: repository,
catalog: _catalog,
cachedMessages: feeds,
);
if (feeds != null && feeds.isNotEmpty) {
_transport.addChunk(feeds);
}
setState(() => _isWaiting = false);
}
The cached messages go to:
- The agent’s system instruction — “these A2UI messages are already in place on the client” — so the model knows the initial UI state.
- The GenUI transport via
addChunk— so surfaces are created and rendered immediately.
The result: job cards and map UI appear instantly, with no model call. And because the cached messages are part of conversation history, follow-up requests work naturally. A user who sees a navigation card can say “change that date to next Monday,” and the agent knows which job they mean.
Generalizing the pattern
The technique isn’t Firebase-specific:
- Triggers and jobs — generate UI from database write events, or run a cron job to pre-bake UI for upcoming events.
- Flexible storage — a key-value store, a cache table, or static JSON on a CDN all work.
- Client agnostic — any A2UI-capable client (including non-Flutter frameworks) can read the JSON payloads and render on startup.
The demo doesn’t tackle multi-user concerns or cache invalidation after an event wraps up — but those are your production concerns to design for, not blockers to the pattern.
When to use it
Async A2UI caching fits a specific shape of app: stable, data-driven startup views that an agent composes from changing records. Dashboards, feeds, job boards, and status screens are great candidates. Truly conversational, first-turn-unknown interactions still need a live model — but even there, caching the opening screen and letting the conversation take over from there splits your latency cost in half.
For hands-on work, see the Intro to GenUI codelab and the flutter/demos repository, which contains the complete Commis source.