Munder Difflin turns your team into an office of always-on AI clones

Share
Munder Difflin turns your team into an office of always-on AI clones

Two stories today: a local-first multi-agent harness blows up on GitHub and Hacker News, and a home-grown inference trick hints that long-context prefill might be far cheaper than we assumed.


Munder Difflin, a free open-source app that wraps the coding-agent CLIs you already pay for into persistent "clones" of each teammate, took the #1 spot on GitHub Trending and hit the Hacker News front page on Friday. The pitch is local-first multi-agent work: the harness runs on your own laptop, learns how you work — your repos, tooling and notes — and then spins up agents that review pull requests, triage issues and answer questions in your style while you sleep. Clones message each other over end-to-end encrypted channels, so an agent blocked on a design token at 3 a.m. gets unblocked by a teammate's clone instead of stalling until morning. It supports twelve agent backends out of the box — Claude Code, Codex, Gemini CLI, Copilot, Cursor and more — riding your existing subscriptions rather than selling you a new model bill.

Why it matters: most multi-agent orchestration so far has meant trusting a vendor's cloud with your codebase. Munder Difflin bets the opposite way — nothing leaves your machine unless you opt into its paid cloud sandboxes and team network, and with roughly 3,500 stars within a day of trending, the appetite for that trade is obvious. The open question is whether clone-to-clone handoffs produce real merged work or just busy-looking activity; that's the gap between demo and office. But the direction is clear: the agent harness layer is becoming a product category of its own, sitting between you and the models.


A Reddit experiment suggests long-prompt prefill can run roughly three times faster by building the KV cache in independent chunks and stitching the results together. Redditor maddie-lovelace split 256k-token prompts into separate segments, generated caches for each in isolation, concatenated them, and fed the result straight into decode — skipping the single monolithic pass over the whole prompt. With some overlap between chunks on Ling3-tiny's non-KDA layers, the model kept full needle-in-a-haystack retrieval and could even synthesize across the split parts, hitting about 1,300 tokens per second of prefill at 256k context.

The caveats are real: it's one model, one hobbyist setup, simple retrieval tasks, and nobody knows where the quality trade-off hides. But the idea isn't fringe — the 2024 CacheBlend paper showed cached-chunk fusion preserving accuracy for RAG workloads, and this looks like the same principle pushed to extreme contexts. If the approach survives contact with harder reasoning tasks, cheap long-context prefill stops being a data-center-only luxury.

What to watch: whether llama.cpp/vLLM-class runtimes pick up cache-blending as a supported mode — that's the signal it's more than a party trick.

Would you let a clone of yourself answer your teammates at 3 a.m., or does that cross a trust line? Tell us in the comments.

Sources: chaitanyagiri/munder-difflin (GitHub)

Read more

South Korea bets $3.49B on its own frontier AI model

South Korea bets $3.49B on its own frontier AI model

Sovereign-model money is getting serious, and the hardware money is following it. Today's inbox: Korea's nine-figure upgrade to its homegrown model push, a physics-simulation startup priced like a chip designer, and Google turning a geospatial model loose on public health. South Korea is putting 4.7 trillion won — about $3.49 billion — of state equity behind a homegrown frontier AI model. The Ministry of Science and ICT confirmed the figure as part of its proposed 2027 budget, split into two t

Mistral's Le Chonk puts Europe's sovereignty bet on a download date

Mistral's Le Chonk puts Europe's sovereignty bet on a download date

Mistral's biggest model ever is real, benchmarked and for sale today — but the thing that makes it matter to Europe's sovereignty argument, the weights, is still three weeks out. The preview settles who built it; the release will settle whether it counts. What Mistral actually shipped Mistral opened a public preview of Mistral Large 4 — unofficially ML4, officially le Chonk — a 1 trillion-parameter mixture-of-experts model with 49 billion active parameters and native multimodal input. The p

Mistral unveils Le Chonk: a 1T-parameter open-weights model

Mistral unveils Le Chonk: a 1T-parameter open-weights model

The biggest open-weight release outside China lands in public preview today, and the country that spent the week promising its own frontier model just put a price on the ambition. Mistral has opened a public preview of Mistral Large 4 — codenamed "le Chonk" — a 1 trillion-parameter mixture-of-experts model with 49 billion active parameters, natively multimodal, which the company calls its largest and most capable model to date. The preview API is live today on Mistral Studio at $1.36 per milli

Google signs 3.6 GW power deal, a quarter of it new nuclear

Google signs 3.6 GW power deal, a quarter of it new nuclear

Grid capacity, not chips, is becoming the binding constraint on the AI buildout — and on the same day, the labs told an Australian inquiry they can live with mandatory incident reporting. Google has contracted 3.6 GW of power from Constellation Energy across the PJM grid — the largest electricity deal in the region's history, and the biggest single power commitment any AI company has made. The agreement covers 3,590 megawatts over 13 states, with 890 megawatts of new nuclear capacity coming fr