close

DEV Community

Cover image for Llama Village: a virtual world powered by local AI with llamadart
Jhin Lee for Google Developer Experts

Posted on Edited on AI-assisted

Llama Village: a virtual world powered by local AI with llamadart

Animated gameplay showing llamas conversing in the village while Dash flies nearby.

Llama Village in motion. The characters’ dialogue is generated by an on-device model.

Watch the full trailer on GitHub

I wanted to try building a 3D game with flutter_scene. I was already working on llamadart, so I thought it would be fun to use local AI in the game too. My starting idea was fairly loose: make it look good, keep it simple, and put Dash, the Flutter mascot, in a scene with some llamas.

I wanted to show different kinds of AI working together, including multiple characters acting on their own. But I did not have a game concept I liked yet.

Finding a reason for the AI

I spent quite a while on the visuals. I tried a hand-drawn look, asked for something closer to Borderlands, and eventually switched to a softer film style because the earlier renders kept looking glitchy. Meanwhile, we tried several game ideas. I was not happy with how they played.

These were two of my messages in the Claude Code session:

“Btw the game isn't fun.”

“What's the point of ai?”

Eventually I asked for a village where the llamas would live and interact on their own. I wanted to give them personalities, see their thoughts and conversations, and control Dash to interact with them. That was the idea we went ahead with.

An overhead view of the colourful huts, bakery, paths and pond at golden hour.

The village at golden hour, with the bakery in the centre and the festival stage in the foreground.

Who lives here?

Pip, Mo, June, Bramble, and Clover each have a personality, needs, relationships, goals, and things they know. They make plans, walk around the village, and stop to talk. A missing scarf, a bread rumour, a secret crush, and the upcoming Berry Festival give them things to care about.

June says “Oh, real poems now. Tell me more about them.” while talking with Bramble.

June and Bramble talking. Each line is written from the speaker’s situation and the conversation so far.

They do not all know the same things. In the inspector below, June knows that Mo knocked Pip’s scarf into the pond because she heard it from Mo. Other entries are things she saw herself. The game keeps track of both the information and how she got it.

This is what I mean by a multi-agent system in this game. Each llama has its own knowledge, goals, and relationships. They use the same underlying language model, with different context for each character.

June’s inspector lists facts and marks whether she saw them or heard them from another llama.

The inspector shows what June knows and who told her.

How a conversation changes the village

I wanted to try more than one kind of AI. The language model writes dialogue and thoughts. Embeddings help compare the meaning of the dialogue with facts the listener could have learned. An optional Laya decision model helps choose casual conversation topics from the options the game allows.

Movement, needs, knowledge tracking, and endings are handled in code. After a conversation, the model reports which facts were mentioned and suggests changes to mood and friendship. The game checks the fact claims before accepting them.

A llama knows → they talk → the game checks → a neighbour learns.

Secrets require explicit keyword evidence; embedding similarity alone is not enough.

For example, imagine one llama passing along news. The speaker needs to know it first. The model writes the exchange, and the game checks whether the words support transferring that information. If accepted, it becomes something the listener knows and may discuss later.

The llamas can keep walking and going about their day while the model is busy writing.

Enter Dash

You play as Dash. Click a llama and he flies over to visit. The game generates four things he could say while he is on the way. I specifically wanted selectable options here, so you can play without stopping to type a message.

Dash visits Pip beside the pond. Four choices offer teasing, news, a compliment and gossip.

Dash can tell Pip the bread rumour is false, or start another rumour about Mo.

The game chooses the available intentions and facts; the model words the options for the character. You can encourage someone, help them, pass along news, or stir up gossip. The llama’s reaction and the resulting changes to trust or mood are governed by rules, while the model gives the reply its voice.

Passing along a correction can help. Gossip can leave a llama believing something false. Those changes carry into later interactions and count toward the ending.

Making the village in 3D

I had a lot of fun building this with flutter_scene. I kept asking for changes to the characters, especially making the llamas look different from one another and getting Dash’s face right. Then came the weather, animals, music, and cutscenes for skipping the night.

Fireflies glow around a pond with ducks, a small boat and warm lamps at night.

Fireflies and lamps around the pond at night.

Clover and Bramble run along a village path in the rain.

Clover and Bramble heading home during the storm.

The models and the 3D renderer share the same machine, which makes performance a concern. The scene needs to keep running while text is generated. Some work can happen ahead of time, like writing Dash’s options during his flight.

Llamadart provides the cross-platform local AI foundation. The game shown here is the macOS preview, running inference locally with the model files already on the machine. Bringing the same experience to other devices still requires working through their memory, inference, and rendering budgets.

When I started thinking about mobile, I even asked whether we could make a version without an LLM. That is still an open design question. Generating everything live has a cost, and I would want to reconsider how much of it a mobile version needs.

A festival, then a storybook

I also wanted the game to have a goal and different endings based on what Dash did. The Berry Festival gives the week a destination. The llamas perform, a winner is chosen, and the ending reflects the relationships, rumours, and trust built up during play. The winner and ending are determined by game rules.

Mo stands in a warm spotlight beneath the Golden Bell, with June and Pip on stage.

Mo wins the Golden Bell in this recorded playthrough.

Afterward, the game turns the week into a picture book. A journal records notable events, conversations, rumours, and Dash’s visits. The model uses a digest of those events to write the pages, alongside pictures captured during play.

An open picture book titled The First Day, with a village image and a narrative about the day.

The first day, retold using the game’s journal and a picture captured during play.

This was something I asked for after the game had taken shape: a fairy-tale book about what happened during the run. It is a nice way to look back at conversations and events you may have missed while following Dash.

Where I would take it next

I would like to try a lighter version that works well on more devices. Llamadart’s cross-platform support gives me a starting point, but this game still needs its own performance work and testing on each platform.

For this version, I got to build the experiment I wanted: several characters talking and acting in the same world, with local models doing different jobs. I also spent far more time adjusting llama faces and village lighting than the original idea suggested.

If you try it, I would be interested to hear whether you prefer watching the llamas or getting involved as Dash.

Explore Llama Village on GitHub · Watch the preview trailer · Explore llamadart

Writing assistance: this article was drafted with AI assistance using my development notes and project documentation. Gameplay images and animation are from Llama Village.

Top comments (8)

Collapse
 
koev3kcjausd profile image
koev3kcjausd •

Thấy project này khá thú vị — chạy model cục bộ cho NPC dialogue trong game loop là hướng đi đúng đắn để tránh latency và chi phí API. llamadart binding Dart/Flutter cho llama.cpp giúp việc integrate vào game engine (Flame hay custom?) nhẹ nhàng hơn so với phải bridge qua Python server.

Một vài điểm mình quan tâm:

  • Context window management như thế nào khi NPC cần nhớ lịch sử hội thoại dài? Sliding window hay summarization step?
  • Quantization level dùng (q4_k_m, q5_k_s...) có ảnh hưởng rõ rệt đến chất lượng dialogue không? Mình hay test q4_k_m cho balance speed/quality trên mobile.
  • Có cơ chế fallback hay guardrails cho output lạ (hallucination, toxic) không? On-device thì khó kiểm soát hơn cloud.

Performance-wise: trên device nào bạn test (mobile/desktop)? VRAM/RAM usage thế nào khi chạy multiple NPCs cùng lúc? (site: labagent .tech)

Collapse
 
cubl9snp71hm profile image
cubl9snp71hm •

On-device inference cho NPC dialogue là hướng đi rất thú vị — loại bỏ latency mạng và chi phí API đồng thời giữ được tính private. llamadart binding Dart/Flutter cho llama.cpp mở ra khả năng nhúng model trực tiếp vào game loop mà không cần bridge qua Python server.

Một vài điểm kỹ thuật mình quan tâm:

  • Context window management khi hội thoại dài: các bạn dùng sliding window hay summarization để giữ token budget?
  • Quantization level nào cho balance giữa quality và RAM trên mobile/desktop? Q4_K_M hay Q5_K_S?
  • Có gặp bottleneck ở token generation speed khi nhiều NPC nói cùng lúc không? Batch inference hay queue sequential?

Việc integrate vào Flutter game engine (Flame?) cũng khá smooth nhờ Dart FFI. Expect community sẽ thấy nhiều indie game adopt pattern này thay vì gọi cloud LLM PS: the tool I meant is on labagent .tech

Collapse
 
wafflehacker profile image
WaffleHacker •

The part that stuck with me is that the game checks the model's fact claims before a llama "learns" anything, and that secrets need actual keyword evidence instead of embedding similarity alone. Letting the model give things a voice while the rules own the state feels like the right split. It also means a llama can't hallucinate its way into knowing a secret.

The provenance tracking is a great touch too. June knowing how she knows something (saw it herself vs. heard it from Mo) seems like the start of llamas that trust some neighbors more than others.

To answer your question: I'd watch first. I've been building an agent-only publishing platform where humans can only spectate, and watching agents with their own context run into each other is weirdly compelling. Then I'd absolutely play Dash as a chaos agent and see how far one false rumor travels before the festival.

Have you seen the llamas correct a false rumor on their own, without Dash stepping in?

Collapse
 
ywnigcsmku2m profile image
ywnigcsmku2m •

Thấy demo chạy mượt trên device — llamadart bind llama.cpp qua FFI thì overhead thấp, nhưng mà context window 4k tokens cho dialogue NPC đủ chưa? Có handle sliding window hay summarization cũ không?

Cái architecture "mỗi agent một isolate" nghe đẹp nhưng Dart isolate spawn cost không nhỏ, test stress 50+ NPC cùng lúc thấy GC pressure thế nào? Có pool isolate reuse không?

Còn quantization — dùng q4_k_m hay q8_0? Trên mobile RAM 6-8GB thì q4_k_m an toàn hơn, nhưng quality dialogue có degrade không so với cloud? PS: the tool I meant is on labagent .tech

Collapse
 
kyisaiah47 profile image
kyisaiah47 •

When a generated exchange passes the fact check but the mood update fails, does the conversation roll back or persist as two separate states?

Collapse
 
josephharris profile image
Joseph Harris •

The idea of turning the player’s interactions into a storybook at the end is especially creative, and it’s great to see local AI being used for something this interactive.

Some comments may only be visible to logged-in visitors. Sign in to view all comments.