Skip to content
All work

NEOSWARM·LOCAL MULTI-AGENT ORCHESTRATOR

Your agents should not need a cloud account to think.

NeoSwarm is a local-first multi-agent system. A team-lead agent breaks a mission into sub-tasks, delegates them to parallel workers, and merges the results. It runs on Ollama by default. Cloud models are opt-in, not assumed.

Role

Sole author

Language

Python, TypeScript, Rust

Runtime

Native Python loop + FastAPI

Providers

Ollama, OpenAI, Anthropic, Gemini

NeoSwarm

Architecture

How the system fits together

Team-lead agent → task decomposition → parallel worker agents → provider registry (Ollama / cloud)

The problem

The demo is not the runtime

Most agent demos are a single prompt to a cloud model, followed by a JSON parse and a screenshot. The real work starts when the network drops, the model returns a different shape, or the task is too large for one context window.

A useful runtime has to plan, delegate, recover, and run where the user works. For a lot of people that means a laptop with no API key budget.

Competitive gap

What existing agent tools get wrong

Single-prompt tools look magical in a tweet but fall apart on multi-step tasks. Frameworks like LangChain help wiring, but they still assume a cloud API key and a single model. None of them treat “run offline” as a first-class constraint.

NeoSwarm is built around the constraint: plan, delegate, and recover locally. Cloud is an opt-in provider, not an assumption.

ApproachWhere it fails
Single LLM callHits context limits and cannot recover from its own errors.
Cloud-only frameworksStop working when the network drops or the API bill runs out.
NeoSwarmLocal by default, multi-agent, model-agnostic, recoverable.

User context

Who this is for

NeoSwarm is for developers and power users who want automation they can run on their own machine without renting a cloud agent. They care about cost, privacy, and being able to debug what happened.

UserNeedWhat NeoSwarm does
Solo developerCheap automationRuns on Ollama; token cost is zero.
Privacy-conscious teamData never leaves the machineLocal loop by default.
Agent researcherSwap models easilyProvider registry abstracts the backend.

What a safe run looks like

The shape of a mission

A mission enters as natural language: "Research the best local LLMs for a MacBook Air, then write a one-page comparison." The team-lead turns that into sub-tasks, each with an output schema and a retry budget. Workers run in parallel where possible. A merger combines outputs that pass validation.

If a worker fails or returns malformed output, it is retried with the schema error attached. If it still fails, the mission surfaces the partial result instead of pretending it succeeded.

The shape of a mission
StageWho decidesWhat can go wrong
Mission intakeTeam-leadAmbiguous scope, missing context.
DecompositionTeam-leadTasks that cannot be merged, circular dependencies.
Worker executionWorker agentsModel errors, malformed JSON, timeouts.
ValidationCodeSchema mismatch, hallucinated fields.
MergeTeam-leadConflicting worker outputs.

Abstraction

The provider registry buys portability

Every provider implements the same three methods: generate, stream, and embed. The registry loads the configured default and falls back if a local model is missing. The rest of the system never imports an SDK directly.

neoswarm/providers/registry.py
class ProviderRegistry:
    def __init__(self, config: ProviderConfig):
        self._provider = self._load(config.default)

    def generate(self, prompt: str, **kwargs) -> str:
        return self._provider.generate(prompt, **kwargs)

    def stream(self, prompt: str, **kwargs):
        yield from self._provider.stream(prompt, **kwargs)

    def embed(self, text: str) -> list[float]:
        return self._provider.embed(text)

    # A model swap is a config change, not a refactor.

Surfaces

Three interfaces for the same runtime

The same Python agent loop powers three frontends. The React/Tauri desktop app is for visual workflows. The Textual terminal UI is for keyboard-driven use. The Python CLI is for scripts and CI.

They all talk to the same FastAPI backend, so a run started in the CLI can be inspected later in the desktop app.

Measured on a MacBook Air M2

What local-first costs and buys

Running Llama 3.2 locally is slower per token than GPT-4o-mini, but the first request has no cold start and there is no bill at the end of the month. The numbers below are from a small benchmark of decomposition + three parallel workers.

~2.4s

First token latency

LLAMA 3.2 3B, QUANTIZED

₹0

Token cost

OLLAMA LOCAL

3

Interfaces

DESKTOP · TERMINAL · CLI

Source · local benchmark, 10 sample missions

Lessons

What building it taught me

The hardest part of multi-agent systems is not parallelism — it is convergence. Four workers producing four answers is easy; four answers that compose into one coherent result is hard. It needs explicit output schemas, retry boundaries, and a team-lead that knows how to say “this does not fit.”

Running locally also forces discipline. You cannot paper over latency with a bigger API budget. The loop has to be small, cacheable, and interruptible.

A multi-agent system is only as good as its refusal to merge garbage.

Still open

Where it is still rough

Convergence is heuristic

The merge step uses structured output and overlap checks, but it cannot guarantee semantic consistency across workers. Some missions need a human merge.

Local models are shallow

Llama 3.2 3B is fast but struggles with complex decomposition. The registry makes swapping easy, but the default experience depends on hardware.

No persistent memory yet

Each mission starts fresh. Long-running context and learned aliases are planned but not implemented.

In short

What NeoSwarm came down to

01

Run offline by default

Local execution is not a feature; it is the constraint that forces the rest of the design to be efficient.

02

Model is a provider

The agent loop should not care whether the model is Ollama, OpenAI, or Anthropic.

03

Merge is the bottleneck

Parallel workers are cheap. Making their outputs compose is the expensive part.

More work