Four workers producing four answers is easy. Getting those answers to compose into something useful is where a multi-agent system starts to become difficult.
That is the main lesson behind NeoSwarm, my local-first agent orchestrator. A team-lead agent decomposes a mission, delegates work, and merges the results. The runtime uses a native Python loop and FastAPI, with a provider registry for local Ollama models and opt-in cloud providers. A React/Tauri desktop app, a Textual terminal interface, and a Python CLI sit around that runtime.
The project is still in progress. The merge step is heuristic, smaller local models struggle with complex planning, and persistent memory is not implemented yet. Those limitations are part of the design problem, so I want to keep them visible while explaining the architecture.
Start with a mission that can fail
Consider the mission used in the case study:
Research the best local LLMs for a MacBook Air, then write a one-page comparison.
There are several jobs inside that request. Someone needs to establish the hardware constraints. Someone needs to identify candidate models. Someone needs to compare them using criteria that make sense for the same machine. The final answer needs to distinguish a recommendation from a result that was measured.
You could put the whole request into one prompt. That might be enough for a short answer, especially if the relevant information is already supplied. Splitting it into workers introduces overhead, and the split needs to earn that overhead.
A multi-agent approach becomes useful when the tasks can be separated without losing the relationships between them. Hardware constraints can inform the candidate list. The comparison depends on both. The final writer needs the same assumptions as the workers whose output it consumes.
Running everything concurrently would ignore those dependencies. A task graph needs to represent them before execution starts.
There is also a practical boundary to the offline promise. A local model can reason over information supplied on the machine, but it cannot fetch current research from the internet while disconnected. Local inference and offline access to source material are separate requirements.
Give workers a contract before giving them a prompt
Natural-language delegation leaves too much room for interpretation. “Find suitable models” could produce a paragraph, a list of names, a hardware recommendation, or a mixture of all three.
NeoSwarm's case study describes sub-tasks with output schemas and retry budgets. The schema gives the runtime something concrete to check before handing an answer to the next stage.
For the model-comparison example, a worker could return a candidate name, the assumptions behind its recommendation, and the evidence it used. A worker that cannot establish a fact should have somewhere to record that uncertainty.
Here is an illustrative output shape, not a claim about the current repository's exact schema:
{
"candidate": "example-model",
"hardware_assumptions": [
"Exact memory capacity must be supplied by the user"
],
"evidence": [],
"unresolved_questions": [
"No throughput measurement is available for this machine"
]
}
An empty evidence list should affect what happens next. The final writer should not turn this record into a confident performance claim because the JSON happened to parse.
Schemas make the expected shape explicit. They do not establish whether a statement is true, whether a source supports it, or whether two workers mean the same thing by “fast.” Those checks belong to a different layer of validation.
Decide which errors deserve a retry
NeoSwarm retries malformed worker output with the schema error attached. If the worker still fails, the mission surfaces the partial result.
That is a more useful recovery path than resending the same request indefinitely. A validation error can tell the worker what to repair: a missing field, an unexpected type, or an invalid structure. A retry budget prevents one task from consuming the entire run.
The budget also forces a decision about what failure means to the rest of the mission. If the hardware-constraints task fails, the comparison may lack a valid basis. If an optional candidate fails, the mission may still produce a narrower comparison with that omission disclosed.
Those failures should not have the same consequence.
A retry also costs work that has already happened. On local hardware, the extra generation competes for the same memory and compute as the other tasks. Retrying every disappointing answer can erase the benefit of splitting the mission in the first place.
I want retries to repair a specific failure. When the failure is an ambiguous task or incompatible assumptions, the plan itself needs attention.
The merge step has to preserve disagreement
The team-lead receives outputs that have passed structural validation. It still has to decide whether they fit together.
In the model-comparison example, imagine one worker recommends a model because it fits within an assumed memory budget. Another recommends a different model based on answer quality. Both may have answered their own sub-task reasonably well. Their conclusions are not interchangeable.
A useful merge needs to retain the criteria behind each answer. It might recommend one candidate for a tighter memory budget and another when quality takes priority. If the machine's memory is unknown, it should leave that dependency unresolved rather than quietly pick an assumption.
A model asked to write a coherent final answer can smooth over these differences. The prose gets cleaner while the reasoning becomes harder to inspect.
That is why convergence is the bottleneck. I need the workers' results to share enough context to compose, and I need the merger to reject combinations that do not make sense.
NeoSwarm currently uses structured output and overlap checks, but the case study explicitly acknowledges that these do not guarantee semantic consistency. Some missions still need a human merge. I would rather expose that limitation than present a polished answer as proof that the workers agreed.
Parallel workers still share one machine
A laptop has a finite memory budget and finite compute. Several logical workers do not necessarily translate into several independent model executions that all finish faster.
The inference backend may queue requests. Concurrent generation may increase memory pressure. A long task may delay shorter tasks even if the Python runtime can schedule them independently. The behavior depends on the model, hardware, and backend configuration.
That leaves two different questions to measure: whether tasks are independent enough to run concurrently, and whether the chosen backend benefits from that concurrency.
The architecture should permit parallel execution where it makes sense without treating worker count as a performance claim. More workers also produce more outputs to validate and more assumptions to reconcile.
For a small mission, the additional coordination may cost more than it saves. For a larger mission with separable work, the split may be worth it. I would want measurements across both kinds of task before making a general claim about speed.
Keep provider SDKs out of the agent loop
NeoSwarm's provider registry exposes generation, streaming, and embedding. The runtime calls that interface rather than importing each provider's SDK throughout the orchestration code.
The boundary keeps provider selection separate from planning and execution. Ollama is the local default. OpenAI, Anthropic, and Gemini are configured cloud options.
The registry pattern can be expressed with a small sketch adapted from the case study. This is illustrative Python; it omits configuration validation and error handling:
class ProviderRegistry:
def __init__(self, provider):
self._provider = provider
def generate(self, prompt: str, **kwargs) -> str:
return self._provider.generate(prompt, **kwargs)
def stream(self, prompt: str, **kwargs):
yield from self._provider.stream(prompt, **kwargs)
def embed(self, text: str) -> list[float]:
return self._provider.embed(text)
A shared interface does not make the providers behaviorally identical. Models differ in how reliably they follow an output schema, how they stream, and how much context they can use. Swapping the provider can change the quality of decomposition as well as the quality of a worker's answer.
Fallback policy also needs to respect the user's choice. Someone choosing local execution may be doing so for privacy, cost, or lack of connectivity. A missing local model should not silently become permission to send their prompt to a cloud service.
The case study describes provider fallback, but it does not establish all of the consent and error-handling details. Those deserve explicit verification before treating the local-first promise as a guarantee for every configuration.
Use the same runtime from different interfaces
NeoSwarm has three interfaces because people approach the work differently.
The desktop app provides a visual surface for workflows. The Textual interface supports terminal use. The CLI supports scripts. They sit around the same Python loop and FastAPI backend.
Keeping the planning and execution rules in that shared runtime reduces the chance that each interface develops its own version of a mission. A retry should mean the same thing from the desktop app as it does from a terminal.
It also concentrates the difficult questions in one place. The runtime needs to determine which work succeeded, which output is usable, and what remains incomplete. Each interface then has to communicate that state clearly.
A completed-looking screen is misleading if a required worker failed. The UI needs to distinguish successful output from a partial result that needs attention. That distinction starts with the runtime's representation of the mission, before any interface renders it.
Local execution has costs worth stating
Local Ollama inference avoids a per-token API charge. It still uses electricity, memory, storage, and the user's time. A model that cannot plan the task well may create more retries or require more human correction.
The case study reports a small local benchmark, but I am leaving its timing figures out of this article. A useful performance discussion needs enough detail to reproduce the workload and understand what was timed.
For this architecture, I would want to separate time to first token from time to a usable final result. The latter includes decomposition, worker execution, validation, retries, and merging. A fast first token does not tell me whether the mission finished correctly.
I would also want to compare the multi-agent run with a simpler single-agent baseline. If the extra coordination does not improve the result or reduce the required human work, the added complexity needs a better justification.
What remains unfinished
The current case study identifies three limits that matter to everyday use.
First, convergence remains heuristic. Valid outputs can still disagree, and overlap checks cannot settle every conflict. Human review remains necessary for some missions.
Second, smaller local models struggle with complex decomposition. Changing the provider can help, but the default experience still depends on the user's machine and chosen model.
Third, there is no persistent memory yet. Each mission starts fresh. Long-running context and learned aliases are planned rather than implemented features.
Persistent memory would introduce more decisions: what to retain, where it came from, how to correct it, and when to stop using it. Adding a store without answering those questions could let an old mistake influence later missions.
The next evaluation I want to define is a set of missions with explicit dependency and disagreement cases. Each run should record whether the merger preserved uncertainty, whether a failed worker changed the final answer, and whether a human could see why the mission was incomplete. That would give me a firmer basis for improving convergence than judging the fluency of the final paragraph.
This article expands the NeoSwarm case study. The project is open source on GitHub. The examples above illustrate design reasoning; they are not a substitute for checking the current implementation.
