Open Source Radar — September 28: the labs open their internal tools
Today's signal is vendors publishing the tools they built for their own engineers — Alibaba's review bot, Tencent's knowledge platform, Cloudflare's audit skill, Anthropic's role plugins — while a one-person project bets that the right answer to Claude Code versus Codex is both.
Open Code Review (Go, Apache-2.0, ~42,000 stars) — Alibaba ran this internally as its official AI code review assistant for two years before incubating it in the open, and the case for that decision is measurable. It reads git diffs, sends changed files to whichever model endpoint you configure, and returns line-level comments, with a built-in multi-language ruleset covering null-pointer, thread-safety, XSS and SQL injection patterns. On its own benchmark — 200 real pull requests from 50 repositories across 10 languages, 1,505 ground-truth issues annotated and cross-checked by more than 80 senior engineers — it reports higher precision and F1 than a general-purpose coding agent on the same model while consuming roughly a ninth of the tokens. Its recall is deliberately lower, which is the trade you want in a reviewer: fewer findings, fewer false alarms. If your team switched off an AI reviewer because the noise cost more than reading the diff yourself, this is the design that answers that complaint.
WeKnora (Go, Tencent, ~30,800 stars) — Tencent's open-source knowledge platform, and the interesting part is what happens to documents after retrieval. Alongside ordinary RAG question answering and a reasoning agent that can call MCP tools and run code in sandboxes, a Wiki Mode has agents distill uploaded material into an interlinked markdown knowledge base with a knowledge graph, revision history and one-click rollback — a wiki that maintains itself from raw sources instead of a pile of embeddings you cannot read. Ingestion covers Feishu, GitLab, Notion, Yuque, DingTalk and RSS, chunk-level editing lets you fix and revert what the retriever actually sees, and long-term memory carries across sessions. It is multi-tenant, ships with more than 20 provider integrations and Langfuse observability, and the licence is Tencent's own rather than a standard one — read it before anything commercial.
security-audit-skill (JavaScript, MIT, ~22,500 stars) — Cloudflare's coding-agent skill for running a real security audit, and the six-phase structure is the story: recon, hunting, validation, reporting, structured output, independent verification, with parallel agents attacking the codebase from different angles. Two phases are the ones most agent-driven audit tools skip. Validation puts separate agents on the job of disproving each finding, and independent verification has fresh agents re-check every factual claim against the actual source. Findings land as schema-validated machine-readable output, and repeated runs are additive because each one reads the previous findings and explores paths nobody has covered yet. This seed grew into Cloudflare's fleet-wide vulnerability harness, so what you get here is the single-repo ancestor of a production system rather than a demo.
Knowledge Work Plugins (anthropics, Apache-2.0, ~25,800 stars) — Eleven plugins, each one a job function: productivity, sales, customer support, product management, marketing, legal, finance, data, enterprise search, bio-research, plus one for building the others. A plugin is nothing but markdown and JSON — skills the model draws on automatically, commands you invoke on purpose, and connectors that wire it into Slack, HubSpot, Snowflake, Figma or PubMed through MCP — so adapting one for your company means editing text, not writing code. Built for Claude Cowork and compatible with Claude Code, and the repository is honest about being a starting point you are meant to rewrite. The caveat is that this is vendor-authored packaging for a vendor product: the value is in the structure and the connector list, and it assumes you are comfortable handing a whole role's workflows to one assistant.
OpenRig (TypeScript, Apache-2.0, ~1,300 stars) — The genuinely fresh entry on today's trending page, and it takes the opposite position to most harness debates: don't choose between Claude Code and Codex, put both in the same team. You declare the rig in a YAML file — which seats exist, what runtime each uses, who reports to whom — boot it, and drive it from a terminal dashboard that shows the seats as a graph, then as a table of runtimes, models, context and state. A lead agent coordinates specialists across teams and returns the decisions that need you, the team keeps its work and context at the same addresses between sessions, and multi-host support spreads one rig across several machines. The caveats are documented rather than buried: it needs Node and tmux, launching a rig writes provider hooks and workspace trust settings on your machine, and at 1,300 stars with a 0.5 version number this is a young project moving fast.
Worth watching this week: whether Open Code Review's precision-over-recall bet holds on your codebase rather than on its benchmark.
Would you rather hand a job function to one broad assistant or split it across narrow specialist agents? Tell us in the comments.
Sources: Open Code Review (GitHub) · Open Code Review · WeKnora (GitHub) · WeKnora · security-audit-skill (GitHub) · Knowledge Work Plugins (GitHub) · Claude Cowork · OpenRig (GitHub) · OpenRig