mission sim gameplay 1 - quochung.cyou PTIT

MissionSim

Inspiration

We keep hearing that AI agents will run factories, coordinate robots, manage infrastructure. Maybe they will. But right now, most multi-agent demos are turn-based puzzles where three chatbots pass notes to each other in a polite loop. Nobody is testing what happens when things go wrong fast and agents have to physically move, fight over scarce resources, and make split-second calls with incomplete information.

That gap bugged us. So we built a simulation where the answer matters: a live facility with a reactor that overheats, oxygen that drains, fuel that runs out, and a crew of LLM-powered agents that either coordinates well enough to survive, or doesn’t.

We wanted something closer to what an autonomous crew on a space station or an offshore rig would actually face. Not a chatbot relay. A pressure test.

mission sim gameplay 1 - quochung.cyou PTIT

What it does

MissionSim is a real-time simulation where a crew of AI agents (Heavy Mechanic, Base Director, Safety Scientist) operates a reactor facility under cascading failures. Each agent runs on its own Qwen model. They see the same world state but have different jobs. They talk to each other in natural language, argue about priorities, and issue commands that move their characters to specific machines on a spatial map.

The reactor generates heat and power. That power has to be split between three systems that all need it at the same time: the Coolant Pump (keeps the reactor from melting down), the Oxygen Generator (keeps the crew breathing), and the Crane Extractor (the only way to win the mission). Oil fuels the reactor, and it runs out. Every allocation decision means something else gets starved.

Gemini Generated Image e82ajde82ajde82a 1 - quochung.cyou PTIT

Here is what makes it more than a demo:

Agents have bodies. The Mechanic at position x=600 cannot allocate power at the Terminal at x=1800 without physically running there. Running costs stamina. If the Mechanic is exhausted, they walk, and the delay can cost the mission. This forces the Director to care about who is closest, not just who is available.

Failures cascade. If the reactor stays above 500°C for 5 seconds, the oil pipe ruptures and drains 50 units of fuel. If someone reallocates power more than 3 times in 10 seconds (panic mode), the coolant pump locks out and needs a physical reset. These failures compound: a pipe rupture means less fuel, which means less power, which means less oxygen, which kills the crew. We watched this happen repeatedly in early testing.

The operator can change the rules mid-run. A judge types “solar surge doubles heat generation” or “cut base power to 50%” in plain language. A Qwen-backed scenario engine translates that into structured game actions (from a set of 18 action types: set heat, trigger failures, lock machines, move NPCs, reallocate power), validates every value against safe ranges, and applies them instantly. No restart, no scripting. The agents see the new alert and have to re-plan on the fly.

mission sim gameplay 2 - quochung.cyou PTIT

Crew size is adjustable. Before each run, you set the squad: 1 to 3 Engineers, exactly 1 Captain, 0 to 3 Scientists. This is how we get a controlled comparison. Same crisis scenario, same physics, different team size. The 3-person crew fails where the 6-person crew survives, and the difference shows up in logged data, not just anecdotes.

Everything is logged. Every speaker-token handoff, every agent decision, every command, every failure and recovery, timestamped and stored in a SQLite database. Past runs can be replayed and compared side by side.

mission sim main - quochung.cyou PTIT

How we built it

The system has two halves that run independently: a deterministic simulation engine (the physics, the machines, the failure triggers) and a non-deterministic agent system (the LLM calls, the decision-making, the coordination).

The frontend is TypeScript and Phaser 4, running a 60fps game loop. The reactor, coolant pump, oxygen generator, crane, terminal, and oil reserve are all entities on a 1D spatial map with defined positions. The physics is frame-based with a scaled delta: during a crisis, the game slows to 0.1x speed. This is called Tactical Dilation, and it exists because LLM response times (1 second polling intervals) are too slow relative to real-time failure cascades. Slowing the game gives agents more “ticks” per crisis-second. The camera always runs at full speed so the viewer experience stays smooth.

The agent system works like this:

A SimulationManager sits between the game loop and the agents. Every tick, it snapshots the world state into a JSON object and pushes it to a MessageBroker. The broker holds the world state, chat history, an action queue, and a speaker-token mutex.

Each agent is an AgentWorker that polls the broker. The Director polls every 500ms; everyone else polls every 1000ms. An agent only calls the Qwen API when there are active system alerts or new messages from other agents, and when it is not busy executing a previous action. Before calling, it acquires a speaker token (a mutex with priority-based preemption, where priority = proximity to the active crisis). This prevents two agents from issuing conflicting commands at the same time.

The prompt going to Qwen has three layers:

  1. A system prompt with the world physics, command JSON schema, and coordination rules. This is the same for all agents and gets interpolated with the actual game constants (machine positions, mechanic rates, failure thresholds) at build time.
  2. A persona prompt specific to the agent’s role. The Mechanic cares about heat and repair. The Director decomposes tasks and delegates. The Scientist monitors oxygen and raises safety concerns.
  3. A dynamic context block: the current WorldState JSON plus the last 10 chat messages formatted as conversation turns.
level1 thumb 1 - quochung.cyou PTIT

The response comes back as JSON with a thought_process, a speak field (what the agent says to the crew), and a command (move, allocate, repair, reset, interact, or standby). If the JSON is malformed, a fallback parser tries extracting it from code fences or raw text. If that fails too, the agent says “I am recalculating my coordinates” and issues a no-op. In hundreds of runs, this has kept the simulation from ever crashing on a bad LLM response.

The scenario engine (ScenarioService) takes a natural-language string from the operator, sends it to Qwen with the current world state, and gets back a structured response with a title, description, and a list of typed actions. Each action is validated against a whitelist of 18 action types and clamped to safe value ranges before execution. Power allocations are ratio-scaled if they exceed the available base power.

The backend is a Python/FastAPI server deployed on Alibaba Cloud ECS via a GitHub Actions CI/CD pipeline. It acts as an LLM proxy with automatic model rotation: if a Qwen model returns 403 or 429, the backend drops that model from the registry and tries the next one. Other errors (401, 500, network failures) do not drop the model. Session data (game records, outcomes, timestamps) is persisted in SQLite.

mission diagram - quochung.cyou PTIT

Challenges we ran into

The first version had all three agents rushing to the Terminal every time a crisis hit. They would queue up, all try to allocate power, and issue conflicting distributions that cancelled each other out. Sometimes the Mechanic would overwrite the Director’s allocation before it even took effect. We solved this with the speaker-token mutex, but tuning the priority function took a few iterations. We landed on spatial proximity: the agent closest to the problem talks first. It is not perfect, but it produces reasonable behavior most of the time.

LLMs do not naturally understand that their character has a body. In early runs, agents would issue allocate commands from across the map. We added a 50px proximity check (the agent must be near the Terminal to allocate), and we rewrote the system prompt to say explicitly: “you must physically move to a machine before you can interact with it.” Even after that, agents occasionally try. The proximity check catches it.

The timing problem was harder. When a failure triggers and resolves within one polling interval (1 second), agents respond to stale state. Tactical Dilation (0.1x game speed during crises) mostly solves this by giving agents 10 effective polling cycles per game-second during the moments that matter most. It is an imperfect solution: the game looks noticeably slower during crises, which is actually a nice side effect because it creates dramatic tension for the viewer.

JSON parsing was a recurring headache. When agents are under pressure (high heat, multiple alerts, long chat history), the Qwen responses sometimes include extra reasoning text outside the JSON block, or wrap the JSON in triple backticks with extra commentary. Our three-tier parser (try code fence extraction, try raw brace extraction, fall back to a safe no-op) has held up, but we still see maybe 5% of responses hit the fallback path in high-stress runs.

Accomplishments that we’re proud of

The Director’s crisis messages are the thing that surprised us most. We did not script any dialogue. The Director prompt says “you decompose tasks and delegate based on agent positions.” What comes out is this:

“Scientist_1, you are 47px from Terminal. MOVE THERE NOW and allocate: Crane=70, Mach_1=20, Mach_2=10. Mechanic_1, move to Mach_1 at x=600 for breaker reset. Mechanic_2, move to Oil Reserve at x=2300 for pipe repair standby. Mechanic_3, standby near Crane.”

That is from an actual run. The Director calculated pixel distances, estimated ETAs, planned a phased power oscillation strategy, and assigned agents to specific machines based on where they currently were. Nobody told it to do that. It emerged from the persona prompt plus the world state.

The fail-versus-win comparison with different crew sizes is the result we are most confident about. Running 1/1/1 (one engineer, one captain, one scientist) against the same crisis that a 3/1/2 crew handles comfortably is a clean, repeatable demonstration. The 3-person crew runs out of hands. The 6-person crew decomposes the same problem and distributes the work. The log data shows the difference in time-to-stabilize and mission outcome.

The scenario injection working in real time was not guaranteed. Translating “solar surge doubles heat” into a set of typed, validated game actions and applying them while the simulation is running, without any downtime, took more plumbing than we expected. But it works, and it means a judge can construct arbitrary stress tests on the spot.

What we learned

The speaker-token mutex was the single most important design decision. Before we had it, agents sabotaged each other constantly. Two agents calling the API at the same time would produce two conflicting allocation commands, and the second one would overwrite the first before anyone could react. With the mutex and priority-based preemption, agents naturally take turns in order of urgency. It is a simple idea, but it changed everything about how the system behaves.

Giving agents bodies changed how they think. When the Director knows the Mechanic is 690px from the Terminal and the Scientist is 47px away, it assigns the Scientist to allocate power and sends the Mechanic somewhere more useful. Without spatial constraints, every agent is interchangeable, and coordination becomes trivial. With them, delegation is a real optimization problem.

Splitting prompts into three layers (system rules, persona, dynamic context) made debugging possible. When an agent does something weird, we can check whether the problem is in the physics rules (layer 1), the persona objectives (layer 2), or the current state representation (layer 3). A single monolithic prompt would be opaque.

What’s next for MissionSim

We want to add more scenario types: a Mars base, an offshore rig, a data center. Same agent architecture, different failure physics.

Model comparison is the next obvious use case. Run the same scenario with Qwen-Plus versus Qwen-Max and compare decision quality from the logged traces. The infrastructure is already there; we just need to build the comparison dashboard.

Agent memory across sessions is the harder problem. Right now, agents start fresh every run. If a crew that failed a pipe rupture could remember what went wrong and try a different approach next time, that would be a meaningful step toward real learning.

We would also like to publish a standardized scenario suite with scoring rubrics, so other teams can test their multi-agent systems against the same challenges and compare results.

Built with

Qwen Cloud, TypeScript, Phaser 4, Vite, Python, FastAPI, SQLite, Docker, Alibaba Cloud ECS, GitHub Actions

Try it out

Video demo:

Published by

Nguyễn Quốc Hưng

I'm delighted to see you here :>

Leave a Reply