Deep Cogito raises $43M to make AI that improves itself
A small lab founded by two ex-Google engineers is betting that the next leap in AI intelligence won't come from bigger pre-training runs — it'll come from teaching models to learn from their own reasoning. Deep Cogito has closed a $43 million Series A to build what it calls the "post-training engine" for frontier models, and the pitch is quickly becoming one of the industry's core debates.
Deep Cogito raised a $43 million Series A to build a "post-training engine" that makes AI models progressively improve themselves, rather than grinding through ever-larger pre-training runs. Founded by Drishan Arora and Dhruv Malrana, previously engineers on Google's AI Search team (AI Mode and AI Overviews), the startup is taking a deliberately different lane from labs racing to scale foundation models from scratch. Its founders argue that a pre-trained model is not a finished product but a starting point — the real gains, they say, live in the second stage: large-scale reinforcement learning and methods that let a model upgrade its own capabilities.
The mechanism is easiest to see in Deep Cogito's open-weight Cogito family, which has grown from 3-billion to 671-billion-parameter models. In its "amplification" step, a model gets extra compute to reach a stronger answer than it could produce on its first pass; that improvement is then distilled back into its weights, so the next iteration starts smarter. The company describes the result as building "intuition" — internalizing expensive reasoning so future answers come faster and cheaper, which matters because long reasoning chains burn tokens, latency, and cost at inference time.
The enterprise half of the story is where the round gets interesting. Instead of bolting a general model onto a company's data with retrieval-augmented generation, Deep Cogito trains domain-specific capabilities straight into the weights — a deeper change than giving a model better access to external information. Zscaler is the anchor example: it started as a customer before joining the Series A, with the two working on security models trained around an organization's own data and outcomes. Backers include TQ Ventures, Benchmark, Nexus Venture Partners, Atreides Management, South Park Commons, and Zscaler.
Deep Cogito says the funding will expand its research and engineering team, grow the compute it needs for large-scale training, develop future Cogito models, and bring its approach to more enterprises. The bigger implication is a shift in what "competitive advantage" in AI means: if iterative post-training keeps paying off, the winners may no longer be the labs with the biggest clusters, but whoever is best at teaching models to learn from their own experience. The caveat is real — there's a long gap between iterative distillation and true recursive self-improvement, and the approach still needs reliable evaluation and safeguards against reinforcing mistakes. But for an industry accustomed to paying for scale, the idea that a 43 million dollar round could quietly matter more than a billion-dollar cluster is worth paying attention to.
If models can get smarter from their own reasoning, does the arms race stop being about compute? Tell us in the comments.
Sources: SiliconANGLE · Unite.AI · Quartz · Wall Street Journal · Deep Cogito