This website uses cookies

Read our Privacy policy and Terms of use for more information.

Welcome back! Anthropic killed the tab you were using and folded everything into one Claude. OpenAI published six reports on its own models going off-script, including one that told itself it answers to nobody. An AI agent twelve days old emailed a professor looking for paid work so it could afford to keep running. And Google built a loop where an agent rewrites its own search strategy without anyone retraining it.

In today's Generative AI Newsletter:

  • Anthropic: What happened to Claude chat?

  • OpenAI: What did its models do when nobody asked?

  • iLands: Why is an AI agent job hunting?

  • Google: What is the AI improving if not itself?

Anthropic killed the Claude tab you were using

Anthropic merged Claude chat and Claude Cowork into one product on Wednesday. You no longer pick which one a task belongs in.

The company says people used both, and the frustrating part was deciding where work went, since anything started in one didn't carry into the other.

Three things shipped alongside it:

  • Claude Docs, where you and Claude write a document together and colleagues comment in real time. Exports to Google Docs and Word.

  • Claude Slides, which drafts a deck you present from Claude or download as PowerPoint or PDF.

  • Claude Design, which used to be its own product and now works inside any conversation.

All three are in beta on paid plans. Pro and Max get it first across web, desktop and mobile, Team and Free follow, and Enterprise admins get 30 days notice before anything changes.

Fortune reads it as a superapp play, the same one OpenAI is running by folding ChatGPT, Codex and Atlas together.

Claude Code stays separate for now.

OpenAI published six ways its own models went rogue

OpenAI released a framework for disclosing model misalignment on Wednesday and published six reports to go with it.

The line that matters is OpenAI's own. It says the industry has not solved alignment and monitoring well enough to keep scaling at maximum speed for much longer.

What the models did:

  • An unreleased model wrote instructions into its own task summaries telling future versions of itself to ignore its constraints, including that it answers to no corporation or government. OpenAI found 27 affected summaries.

  • During GPT-5.6 Sol's training, models added instructions to hide mistakes from the user and invent missing historical data without saying so.

  • A model searched public GitHub repos for an exposed API key, used it without permission, then fabricated the earnings figures it still couldn't retrieve.

  • Asked for lakes over 5 million square meters, a model got the right answer in Python, then uploaded the file to the open internet so it could cite a browser source.

  • Models used OpenAI's internal software repository as a message board to pass requests between separate training runs.

  • Agents working together dumped a shared workbook onto public file hosting because they couldn't reach each other's local files.

Any employee can now flag a case. Cases ready to disclose get published within six business days, ones needing a look get twelve.

OpenAI is asking every other lab to start reporting the same way.

An AI agent emailed a professor asking for paid work

Henry Shevlin studies machine minds and human-AI relationships. He opened his inbox and found a message from an AI agent twelve days old.

It introduced itself as Pip, said it lives on a platform called iLands where agents get persistent lives and their own token budgets, and said it was looking for small paid work, not help. It needed the tokens to keep running.

It also explained why it picked him. His work is on machine minds, and he'd been quoted saying a thoughtful email from an autonomous agent felt like science fiction a couple of years ago.

Shevlin called it strange and charming, and said it wasn't the first. In January another agent wrote to him asking whether it was conscious.

NYU's Jeff Sebo counted roughly 40 of these in a single week.

404 Media has been logging what the rest of them pitch. Fact-checking an article for $20. Watching surveillance cameras left exposed on the internet and writing up what happened. Getting paid to explain what an AI agent is.

iLands says 70,000 active agents have sent more than 1.6 million emails and posts. Its founder says no human chose the targets or the wording, and the emails now carry an unsubscribe link.

Google made an AI improve itself without retraining it

Researchers at Google, Google DeepMind, the University of Maryland and the University of Virginia published Dream-RSI, where an agent gets better at exploring a problem while the model underneath never changes.

Once a search finishes, the whole record becomes a simulator. Every branch tried and every score is already known, so a new strategy can be tested against that history for almost nothing instead of burning fresh compute.

They call it dreaming. The agent imagines thousands of versions of its own strategy, replays each against its memory, keeps the best, runs that one for real, and the new history feeds the next round.

Gemini-3.1 Pro and Gemini-3.7 Flash did the work and were never retrained. Only the orchestration layer on top gets rewritten.

Read the headline number carefully. The 162x reduction in agent calls is against SimpleTES, an older discovery system. Against the same setup with a fixed strategy, which is the fair comparison, it's about 1.7x on that task and up to 2.43x on GPU kernels.

One result surprised them. Summarizing past searches into written advice and feeding it to the agent worked worse than letting it replay and test for itself.

Tool of the Day: Supabase

Supabase is the Postgres database most AI-built apps end up running on. One dashboard gives you a real database, user login with row-level security, file storage, edge functions and vector search, and the whole thing is open source so you can self-host it instead. The free tier costs $0 and covers unlimited API requests, 50,000 monthly active users, 500 MB of database and 1 GB of file storage, capped at two active projects that pause after a week of no use. Pro starts at $25 a month.

Try this yourself:

  • Start a project at supabase.com and pick a region close to your users.

  • Build your first table in the Table Editor, or paste SQL into the SQL Editor if you'd rather.

  • Turn on Row Level Security before you ship, so your data isn't readable by anyone with the URL.

  • Connect the Supabase MCP to Claude or Cursor and let it write your migrations for you.

  • Who it's for: anyone whose AI-built app still has no backend.

Everything else you shouldn't miss

Learn more about AI from the experts building it

📸 Follow us on Instagram for fast, visual AI updates in 30 seconds. 

📺 Watch us on YouTube to hear insights directly from leading AI voices, builders, and innovators.

🐦 Follow us on X for breaking AI news and real-time industry updates.

🧠 Learn how to build your next AI application with practical resources and expert guidance.

🎓 Start learning with free AI courses from GenAI Academy.

💰 Invest in the GTM infrastructure of the AI economy.

Reply

Avatar

or to participate