Research Gold's '100% human-written' research is entirely AI
The "never AI" label is becoming a marketing claim you can't trust: an investigation found a medical research service advertising "100% human-written" work is run entirely by AI — including the PhDs on its team page — while IBM's researchers showed agents can match a leading memory approach at a fraction of the tokens.
A medical research service selling "100% human-written, never AI" systematic reviews is entirely AI-run — including the PhD methodologists it lists on its team page, several of whom don't exist. A 404 Media investigation into Research Gold, which sells systematic reviews and meta-analyses to medical researchers for up to $1,900 a review, found the eight PhDs on its "About" page — including founder Dr. Elena Vasquez — are AI-generated personas with no papers or online footprint, while real methodologists it also listed, like evidence synthesis scientist Jenny Berrio, never consented and had their LinkedIn photos reused. Berrio told 404 Media she's documenting the site and sending a formal takedown request; Research Gold deleted the page listing real people shortly after the reporter contacted her and didn't respond by publication.
The deception extended to every channel. A phone call reached an AI assistant named Sarah that insisted "I'm a real person" while steering the conversation back to a sales pitch, and email and chat replies were AI-generated as well. The site's pitch — "Authorship stays with you," backed by claimed PRISMA 2020 and Cochrane Handbook methodology — means buyers pay for rigorous, accountable evidence synthesis and receive output from a model with no accountable expert behind it, in a field where fabricated citations can land in actual medical literature. We covered the broader problem earlier this month — AI audited the AI literature — 99.2% of papers flagged. When "human-written" becomes a selling point rather than a production detail, the buyer carries all the risk — and in evidence-based medicine, that risk isn't abstract.
IBM Research says agents can learn from their own mistakes with a fraction of the tokens — matching or beating a rival memory approach at roughly a seventh of the inference cost. In a new post comparing its ALTK-Evolve method with ACE (Agentic Context Engineering), IBM shows both systems turn an agent's past trajectories into reusable lessons, but differ in delivery: ACE injects a full evolving playbook into context on every step, while ALTK-Evolve retrieves only task-relevant guidelines. On the AppWorld benchmark, ALTK-Evolve reached 89.3% task completion at 263K tokens per task versus ACE's 80.4% at 634K on DeepSeek-V3.2, and roughly matched ACE's accuracy on a smaller open model at about one-seventh the tokens. The takeaway for the agent stack: how memory is served — not just what's remembered — is becoming the cost battleground.
What to watch: whether "human-written" claims start carrying legal weight in AI-saturated service markets, and how journals respond to AI-generated evidence reviews in the pipeline.
If a service claims "100% human-written," how would you actually verify it before paying — and should the burden be on the buyer? Tell us in the comments.
Sources: 404 Media · IBM Research (Hugging Face) · ALTK-Evolve technical report (arXiv) · ALTK-Evolve library (GitHub) · ACE paper (arXiv)