Open Source Radar — October 8: The Agent Wranglers

Share
Open Source Radar — October 8: The Agent Wranglers

Today's trending list is less about new frameworks and more about the awkward middle of agent development: supervising fleets of coding agents, auditing what they touch, and keeping them fed on long tasks. Four fresh projects, all verified on GitHub.

cmux (Swift, ~28,000 stars) is an open-source macOS terminal built specifically for running AI coding agents — Ghostty-based, with vertical tabs, per-agent notifications, and everything exposed through a CLI and Unix socket so agents can drive the terminal themselves. It's climbing daily trending because a terminal designed around parallel agent sessions is suddenly a normal thing to want. Use it as the cockpit when you run more than one coding agent at a time.

security-audit-skill (JavaScript) is Cloudflare's open-source coding-agent skill that turns your agent into a multi-phase security auditor: reconnaissance, coverage-led hunting, validation, then machine-readable findings that get independently checked before they ship. Cloudflare says this skill seeded the internal vulnerability-discovery harness it now runs fleet-wide, so this is the single-repo starting point of a production system. Grab it if you've ever wondered whether your agent can actually review code for security holes rather than just style.

hallmonitor (Go) is a read-only dashboard for your coding-agent fleet: every Claude Code and Codex session — on your laptop and your SSH boxes — shown live in a terminal board, menu bar, or right around the MacBook notch, with who's working, who's stuck, and what each one is doing. It launched on Product Hunt today, and the privacy stance is the selling point: it never sends prompts or approves anything, it just watches. Reach for it when you've lost track of which terminal that agent was in.

pmtui (Rust) tackles the longer-horizon version of the same problem: a tmux-backed supervisor that keeps one persistent Claude or Codex session per project nudged toward a goal you set, auto-handles the ordinary choices, and only pulls you in when a decision genuinely needs a human. The insight is that agent loops drift on long tasks, so pmtui keeps the loop outside the agent and rebuilds the prompt from your goal plus the agent's own last status each time. It hit Product Hunt today with its first releases this week — a bet on supervision layers over agent runtimes.

Worth watching this week.

Supervising agents is becoming its own tooling category — which of these would survive a real week of your workflows? Tell us in the comments. Sources: cmux (GitHub) · security-audit-skill (GitHub) · hallmonitor (GitHub) · pmtui (GitHub)

Read more

Deep Dive — The four-token blind spot inside DeepSeek V4

Deep Dive — The four-token blind spot inside DeepSeek V4

ByteDance's Seed research team says it has found the cause of one of the stranger recurring complaints about DeepSeek's models: the same question, asked with nothing changed except a few junk characters bolted onto the front, can flip the model from right to wrong. Their paper, posted to arXiv on September 28, traces the wobble to a memory-saving trick used during long-context inference, and reports that DeepSeek-V4-Flash-Base's retrieval accuracy swings by as much as 40.2 percentage points depe

SoftBank seeks $100B from Gulf investors for an AI fund

SoftBank seeks $100B from Gulf investors for an AI fund

Three moves today point the same direction: the money, the politics, and the price of speed all got more expensive. SoftBank is reportedly seeking up to $100 billion from Gulf investors for a fund that would buy companies and run them with AI. The Financial Times reported the raise, citing people familiar with the matter, and says Masayoshi Son has held discussions in recent weeks with senior figures including in the United Arab Emirates; Reuters and Bloomberg both carried the report but neith

AI 101 — What is a jailbreak?

AI 101 — What is a jailbreak?

A jailbreak is a prompt — or a carefully arranged stack of inputs — engineered to talk an AI system past its own safety rules, so that it produces content or takes actions it would normally refuse. Nothing is broken in the technical sense: the model, the servers, and the locks all keep working. What gets broken is the instruction to say no. Why it matters right now The word "jailbreak" shows up constantly in AI coverage — in stories about chatbots misbehaving, about guardrails, about agents

Samsung projects a 100 trillion won quarter on AI memory demand

Samsung projects a 100 trillion won quarter on AI memory demand

The AI buildout's money keeps landing in the same place — memory — and Samsung just put the biggest number yet on it. Meanwhile, China's leading open-model lab walked through what its next models still can't do. Samsung projected third-quarter operating profit of 107.4 trillion won — roughly $80 billion — which would be the first time any company has cleared 100 trillion won in a single quarter. The preliminary guidance, released Thursday in Seoul, compares with 12.17 trillion won a year ago,