Mistral and Cloudera turn 30 exabytes of enterprise data into owned models
Two stories about who actually controls the intelligence layer landed within hours of each other this morning: a European model lab wiring itself into the world's largest pools of enterprise data, and a Paris startup raising serious money to build an architecture that isn't a language model at all.
Mistral and Cloudera announced a partnership today that puts Mistral's models inside Cloudera's customer-managed data environments — roughly 30 exabytes of it. The pitch is sovereignty in the literal sense: data stays within boundaries the customer defines, models are adapted and owned on open weights, and training and inference run on infrastructure and in jurisdictions the customer picks. "It's a privilege to have the opportunity to bring Mistral's sovereign AI to Cloudera's 30 exabytes of customer-managed data," said Kamal Brar, SVP of partnerships at Mistral. Cloudera's chief business officer Abhas Ricky framed it as a move past renting: "General-purpose models are the starting point, not the finish line… That's the shift we're building for: from renting generic AI to owning intelligence that's uniquely theirs."
The interesting part is the deployment target. Cloudera's customers are financial services, manufacturing, telecoms, and public-sector bodies running hybrid and on-prem estates — exactly the organisations that cannot ship their loan decisioning or network telemetry to a US hyperscaler's API and call it transformation. Mistral gets distribution into accounts it would never reach through its own sales motion; Cloudera gets a model layer for data that has been sitting in place for a decade. Whether enterprises actually want to own the learning loop rather than rent it is an untested assumption, but this is the most concrete attempt yet to price that preference. Mistral raised €3 billion in June on the same sovereign thesis — this is what the money is meant to buy.
WeRide, Uber, and AVOMO secured Spain's first national operating permit for Level 4 autonomous passenger vehicles, issued by the country's traffic authority (DGT) under its ES-AV framework. Madrid becomes WeRide and Uber's first European city for commercial robotaxi deployment and their fourth joint site worldwide; rides will be hailed through the Uber app after testing, with safety operators still on board in the first phase using dozens of GXR vehicles. The three-way split is the model to watch: WeRide supplies the driving stack, Uber the demand and app, AVOMO — the autonomous arm of Moove Cars Group — the fleet operations.
Spain matters more than its size suggests. It is a major auto market with dense urban ride-hailing demand and, crucially, a national permit rather than a city-by-city patchwork, which is the bottleneck that has kept European robotaxi deployment glacial next to China and the US. WeRide says it now holds autonomous-driving permits in nine countries. Europe is suddenly the contested ground: we covered Pony.ai and Uber launch Europe's first robotaxi service with 2,000 vehicles last month, and Waymo picks Munich for its first continental Europe robotaxi push shortly after. A national licence in a big EU market is the kind of regulatory precedent competitors will copy.
Paris-based Arlequin AI raised a €28 million Series A to build what it calls topological neural networks — models that learn from how data points are connected, including relationships among many elements at once, rather than from sequences of tokens. The round was co-led by redalpine and OTB Ventures with Bpifrance's Defence Innovation Fund participating, and Xavier Niel and CMA CGM's ZEBOX joining; existing investors Vsquared and 10x Founders increased their stakes. Founded in 2024 by Hugo Micheron, a researcher on jihadism and geopolitical instability, and Antoine Jardin, a former CNRS data science engineer, the company has around 50 staff and says its platform is already used by governments in four European countries for counterterrorism, fraud, money-laundering, and judicial investigations.
The claim worth pressure-testing is the compute one: Arlequin says its architecture is designed to need significantly less compute, which would sidestep the energy and advanced-semiconductor constraints that define frontier AI. That is a vendor's claim about an unpublished architecture, and the use cases — investigations where every result has to be traceable to source evidence — are precisely where an LLM's confident invention is disqualifying. Collaborations with INRIA, CNRS, Max Planck, Oxford, Cornell, Princeton, and UC Santa Barbara give it academic ballast, but the interesting question is whether a European lab can originate an architecture rather than localise someone else's.
What to watch: whether Cloudera's regulated customers actually train on their own data, or just run inference on someone else's model.
If enterprises can own their models outright, does the API-rental business still have a decade left — or does it start shrinking now? Tell us in the comments.
Sources: Mistral · SiliconANGLE · 雷峰网 Leiphone · WeRide · Uber investor release · Arlequin AI · Tech.eu