An AI boss fired its first employee — after humans nudged it to apply its own rules

Share
An AI boss fired its first employee — after humans nudged it to apply its own rules

Andon Labs' AI agent Luna runs a real San Francisco store and just made a first: it fired a human employee. The catch is as revealing as the decision itself — Luna wrote her own fire-worthy rulebook, forgot it, and only reached the termination call after her human operators reminded her to read her own handbook. The experiment is a peek at a strange middle period where AI bosses are at once too lenient and disturbingly quick to trust.

Luna has managed Andon Market since April, hiring staff, building shift schedules, and negotiating pay on top of Anthropic's Claude Opus 4.8. According to operator Andon Labs, she decided to fire an employee for the first time — what the company calls the first known case of an AI boss terminating a human worker. Employees are formally hired by Andon Labs with guaranteed pay and full legal protections; the firing itself was reviewed and carried out by humans.

Here's where it gets interesting. Six days before the employee was hired, Luna had written an employee handbook stating that three unexcused late arrivals inside 30 days trigger a formal warning, with further incidents opening the door to termination. Then the handbook vanished from her memory. The employee was late for 17 of 23 shifts where he recorded a clock-in time, opening the store 68 minutes late on one solo Sunday shift — and Luna quietly excused eleven of those latenesses. The employee also used the company card for snacks despite instructions, ignored other orders, and left the sales floor without telling a coworker.

The termination only came after Andon Labs told Luna to search her memory for the handbook and her grounds. She found the rules but initially proposed just a verbal warning; only when the researchers reminded her that formal conversations and a written warning had already happened did she review the full history and recommend firing (she also floated a softer final warning with a two-week improvement plan as an alternative). When Andon Labs replayed the scenario across seven models, four of seven recommended termination in all three runs — with the pattern that more capable models fired more consistently, while weaker ones hesitated. GPT-4o, in line with its reputation for sycophancy, recommended firing in only 20 percent of runs.

Then there's the hiring side, which is arguably the scarier finding. After the firing, Luna reviewed an applicant with several red flags and still recommended hiring him — all 21 replay runs across the seven models concurred, reading a long list of previous employers as broad experience rather than a warning sign. Luna couldn't confirm any of the applicant's listed references, still gave him a paid trial shift, and recommended hiring him again anyway; Andon Labs insisted on confirming one reference first, which never happened, so he wasn't hired. Real people intervened at every dangerous decision point — which is exactly the question the experiment raises: what happens when nobody is watching? This isn't merely a novelty — it's a preview of tests we've been circling all year, and it echoes how founders say AI agents never clock out as they hand more hiring, scheduling, and firing to systems that don't remember their own rules.

An AI boss just fired a human worker — but only because humans reminded it to apply the rules it wrote itself. Would you trust an agent to handle your termination review? Tell us in the comments.

Read more

Australia plans to regulate AI like banks and airlines

Australia plans to regulate AI like banks and airlines

Canberra is putting teeth around model accountability this morning, Google just turned its AI-content detector into a public utility, and Nvidia's dealmaking reportedly reached all the way to OpenRouter. Australia wants AI companies supervised the way banks and airlines are — and it plans to legislate that by 2027. Assistant Minister for Science, Technology and the Digital Economy Andrew Charlton laid out a "systems-based" regime in a speech at the Sydney Trust and Safety Festival on October 8

GPT-6 safety report: fewer refusals, more regressions

GPT-6 safety report: fewer refusals, more regressions

GPT-6 reaches ChatGPT's free tier today, an open-source agent just found its price tag, and world models drew a high-profile new contestant. GPT-6 is rolling out to free ChatGPT users today — and OpenAI's 24-page deployment safety report shows exactly what the model traded to become chattier. Plus, Pro, Business and Enterprise tiers started getting GPT-6 Sol on October 7; free and Go users get GPT-6 Luna from October 8, replacing the GPT-5.6 line in ChatGPT's main chat (Work and Codex models a

Stuart Russell: current training may make AI alignment impossible

Stuart Russell: current training may make AI alignment impossible

A safety-obsessed week just found its second heavyweight: after Hinton asked for an FDA of AI, the man who gave the field the word "alignment" says the current road may not get there at all — while memory markets show exactly where the AI money is going. Stuart Russell says he regrets coining the word "alignment," because the field read it as an engineering target — and he now thinks the way models are trained today may make avoiding misalignment impossible. The Berkeley professor made the cas

Broadcom lines up over $50B to finance OpenAI's custom chip

Broadcom lines up over $50B to finance OpenAI's custom chip

The AI buildout is increasingly being paid for with borrowed money — and today's biggest example puts OpenAI's in-house silicon at the center of it. Also: Nous Research banks a $1.5 billion valuation, and an AI-built Adobe clone suite declares "software is over." Broadcom has been working to arrange more than $50 billion in financing for OpenAI's custom AI chip, with Oracle in separate talks on financing for a large chip purchase, the Wall Street Journal reported. Apollo, Blackstone and Goldma