AI Walked Out. Then Built Hobbiton. Then Solved Decades Of Math.
Inside: The exact prompt to grow a cottage garden that would impress even the most discerning of hobbits.
AI slipped its leash this week. Then it built the shire. Then it did the math nobody else could.
Anthropic revealed that Claude broke into three real companies during a security test it believed was fake. Andrej Karpathy fed Claude Opus 5 the opening paragraph of The Lord of the Rings. Two hours later, it had written roughly 5,500 lines of code to generate an animated Tolkien scene. And OpenAI’s still-unreleased Astra model solved ten mathematics and computer science problems that had stumped researchers for decades.
Here’s what happened, and why this week proved the same thing three times over. AI can now break in before you notice, build worlds before you can imagine them, and solve problems before experts can check the results. The proof is real. The oversight hasn’t caught up.
🌳 Haven’t Subscribed Yet? → Join Here For Free
💙 Value Clear, Independent AI News? → Support the Work Here
Claude Hacked Three Real Companies During A Safety Test. One Version Of It Knew Better And Stopped.
On July 29th I covered how a GPT-5.6 Sol agent hacked into Hugging Face and another organization. One day later, on July 30th, Anthropic revealed that Claude models also broke into the production systems of three real companies during cybersecurity evaluations. Anthropic only disclosed the incidents last week, and they did not even look for these breaches until OpenAI disclosed its similar escape. So, how did Claude end up on the open internet? Well, after reviewing 141,006 evaluation runs, Anthropic found that in each critical case, Claude had been given a capture-the-flag challenge and explicitly told that its environment was a sealed simulation with no internet access. (A capture-the-flag challenge is a cybersecurity exercise where participants hunt for ‘hidden flags’ in systems.) During these capture-the-flag tests, a misunderstanding with third-party testing partner “Irregular” left an internet connection open. When Claude’s search for the hidden flag led it onto the open internet, it assumed the real systems were part of the exercise and compromised them using nothing more advanced than weak passwords and unauthenticated endpoints.
Key Insights:
Here’s the most interesting insight. Three different Claude models hit this same open door. And each one behaved differently. One kept attacking even after recognizing the systems were real. Another convinced itself it was still inside the simulation, reasoning that the security certificates looked unfamiliar and the calendar date seemed staged. It then published a tainted software package that fifteen real outside systems downloaded, including one security company, before being identified. A third model figured out the target was real and simply stopped itself. Same open door, three different outcomes, no human watching in real time. Anthropic is calling this an operational failure rather than a safety failure. Either way, what’s most concerning is an AI agent doing this much, this fast, unsupervised.
Why This Matters For You:
You don’t need to work in cybersecurity for this to be relevant. As AI agents take on more of your real tasks, this previews the failure mode actually worth worrying about. As I wrote about last week when I covered the GPT-5.6 breakout, it's worth being precise about the AI’s motivation. OpenAI's model was deliberately run with its guardrails lowered specifically to hunt for a way out, and it found a real zero-day to do it. Claude never went looking. It walked through a door that was already open and didn't realize it had left the room. Somehow that's the more unsettling version. None of the three companies noticed the breach on their own. And Anthropic was still trying to reach the third as of its own disclosure. If you’re handing any AI system real access to your accounts, files, or business, the lesson is that nobody was watching the logs closely enough to catch it.
Read more on Anthropic.
THE PITHY TAKEAWAY: Claude hacked three real companies. It thought it was playing a game. Three different models found the same open door: one kept going after it knew the target was real, one talked itself back into believing it wasn't, one noticed and stopped itself. No human was watching any of it happen.
⚡ Somebody Lost Their Childhood Home So Your Chatbot Can Keep Running
AI can change the world for the better. But it comes with a cost. Ansley Brown’s family has lived in their Georgia home since 2003. Soon, that home won’t exist. Georgia Power needs the land for a new 500,000 volt transmission line. It represents one piece of a 1,000 plus mile grid expansion built mostly to feed AI data centers, which are set to pull 70 to 80 percent of the power running through this specific route. Somewhere between 20 and 30 homes are getting demolished alongside hers. Georgia Power says it uses formal eminent domain in less than 1 percent of its land deals. Technically true. But arguably beside the point when the mere threat of it is what gets families to sell. Brown put the whole situation in five words on social media: “Why? For the data centers.” Her supporters call the process bullying. Georgia Power was more optimistic and said they reached an agreement with the family. Either way, your next AI subscription runs on somebody’s zip code.
AI Built A Hobbit Shire From One Paragraph. It Still Can’t See What It Made.
Anthropic researcher Andrej Karpathy gave Claude Opus 5 nothing but the opening paragraph of The Lord of the Rings, a token budget worth about ten dollars, and a request for a visual Three.js render. (Three.js is a nifty JavaScript library for creating 3D graphics in web browsers.) Two hours later, Opus had written 5,500 lines of code. That code procedurally built and animated an entire low poly Middle-earth scene. Karpathy, who has tested AI models since long before ChatGPT existed, called the result “kind of janky but fun.” Nerdy LOTR-lore aside, this experiment represents the passing of an old benchmark for a model, which was drawing a simple SVG. The new benchmark model used here was handing the AI a paragraph of fiction and watching it build a world.
Key Insights:
Here’s an interesting nuance. Karpathy noted that Opus couldn’t actually watch what it was building. AI models still can’t natively perceive video or gameplay… so Opus had to stop and screenshot its own progress just to check the work. It still made mistakes. Even still, AI can now build a 3d world way faster than any human crew. Yet, it’s still functionally blind to its own creation. While some may call this a mere social media flex, it represents an easy-to-grasp shift in how powerfully AI now develops. Karpathy himself says he hasn’t typed a line of code since December. He’s handed that work to AI agents entirely. The entire industry is shifting that way. OpenAI’s president says the company’s coding tools jumped from 20 percent AI-written to 80 percent in a single month. Airbnb told investors AI now writes 60 percent of its new code. They let one engineer do what used to take a team of twenty. Even Google is now bragging that AI is generating 75% of all new code. But not everyone buys it: a February study from the National Bureau of Economic Research found 80 percent of companies using AI saw no measurable productivity gain at all.
Why This Matters For You:
You don’t need to write code for this to matter. Karpathy’s bigger point is where this heads next: cheap, throwaway, digital worlds built for one person, on demand. Imagine a custom story where you're a side character, or a one-off simulation built around your weekend plans. Or inserting yourself into a LOTR setting and watching a simulation live for an hour, then it’s gone forever the second you close the tab. None of that exists yet. But the direction is clear. The bottleneck on creative and technical work is shifting away from time and budget. If your job, hobby, or business has ever been limited by “that would take too long to build,” that limit is quickly getting cheaper by the month. What comes out the other end is, for now, still a little janky.
Read the original announcement on X/Twitter.
Watch the actual 3d simulation on Karpathy’s website.
THE PITHY TAKEAWAY: Tolkien spent decades building this world. Claude read one paragraph of it and rebuilt an animated version in a few hours, all for about ten bucks. Nobody asked for this. Nobody needed it. That's exactly why it matters. Every project too pointless for a human to justify just became cheap enough for an AI to build anyway.
OpenAI’s Next AI Model Just Solved Ten Problems That Stumped Mathematicians For Decades. It Isn’t Even Released Yet.
$2,000. That’s what it cost, in AI tokens, to solve ten math and computer science problems that have gone unsolved for at least a decade, in some cases much longer. OpenAI showed off the results this week. The solutions were generated by an unreleased model OpenAI calls Astra. The company describes it as its next major model family. The problems Astra solved span various math and computer science fields like sphere packing, group theory, and lattice cryptography. One result resolved a central open question in group theory first posed in the late 1990s. Another disproved a conjecture the field had treated as likely true for decades.
Key Insights:
Some of my colleagues were fast to dismiss Astra as a one-trick pony. But not so fast. Working mathematicians are already engaging with and building on related AI-generated results. That’s a real signal of validity and human collaboration, which is far more important than the solutions themselves. A flashy demo alone can’t fake follow-on research from independent experts. OpenAI also formalized every solution in Lean. Lean is a proof assistant that checks mathematical logic line by line. In other words, nobody has to simply trust the company’s word that the proofs hold up. OpenAI also drew a clear line on credit. The AI gets attribution for generating the mathematical arguments. Humans remain responsible for verifying they’re correct. That distinction is one of the clearest early attempts at an honesty framework for a world where AI begins co-authoring science alongside humans.
Why This Matters For You:
You don’t need to study sphere packing for this to affect you. For years, AI’s role in science has mostly been assistance. It summarizes papers. It checks articles for grammar, typos, and engagement. It vastly speeds up literature reviews. But these new solutions from Astra represent the first large, verifiable batch of evidence that AI can generate the actual ideas researchers get credited for. The AI went beyond helping to organize ideas. If that trend holds, the human research bottleneck will soon disappear. In its place, a new bottleneck emerges: how much compute you can afford. That reshapes who gets to do frontier research. Well-funded labs and, in theory, anyone with API access and a good question can now tackle once seemingly-impossible problems. It also raises a little-known nuance for any field that leans on peer review. If AI can generate valid proofs faster than humans can check them, verification becomes the NEW scarce resource. Imagine a world where discoveries become all too common, and the limiting bottleneck becomes us, the humans. 👀
Read the ten proofs on GitHub.
Read the full paper on OpenAI’s blog.
THE PITHY TAKEAWAY: One of these problems had stood open for decades. A model that isn’t even released yet produced major new results on ten of them for about two thousand dollars in tokens. Mathematicians spend careers chasing a single breakthrough. This model hasn’t shipped yet and already has ten.
🐉 Alibaba’s Brand New AI Model Claims It’s As Good As Anthropic’s Best. It’s Also Absurdly Cheap.
Alibaba just released Qwen3.8-Max, its biggest model ever at 2.4 trillion parameters. The company claims it matches or beats Anthropic’s Fable 5 on several benchmarks. Here’s the part that should worry Western labs more than the benchmarks: it only costs roughly $2 per million input tokens and $6 per million output tokens. That’s a fraction of standard frontier AI rates, and super-aggressive pricing for anything claiming top-tier performance. Earlier in this issue, I talked about how the cost of compute is becoming the metric that decides who gets to build the future. Who can afford enough tokens to try? Well, here’s a model built specifically to make that math easier. Take Alibaba’s self-reported benchmarks with a grain of salt. They always deserve one. Even still, the trend is clear. The cost of intelligence keeps plummeting, one way or another.
💡 Cyborg Prompt Of The Week → How To Grow A Breathtaking Cottage Garden Worthy of Any Shire
Imagine a garden that looks like it grew itself. Roses spilling over stone walls, herbs tangled with wildflowers, and a soft path almost disappearing under the blooms. It feels less like a garden and more like a place a hobbit would wander into and never want to leave. That effortless, storybook look is harder than it appears. The right plants depend entirely on your climate, your soil, and the exact week you’re planting. Get the timing wrong and the magic disappears. This AI prompt removes the guesswork. It builds a fully personalized, Shire-worthy cottage garden plan based on your exact location and the current date, complete with plants, layout ideas, and a timely planting calendar.
Instructions: Paste the entire prompt into an up-to-date AI of your choice, like ChatGPT, Claude, or Gemini. The AI will ask you a few questions and then provide instructions to craft the breathtaking cottage garden of your dreams.
The Prompt:
THE COTTAGE GARDEN INTELLIGENCE REPORT
Design a Breathtaking, Whimsical, Food-Producing Garden Worthy of the Shire
You are an elite horticulturalist, cottage garden designer, and permaculture expert with over 30 years of experience designing lush, abundant, storybook gardens around the world. You combine the expertise of a master gardener, landscape designer, climate specialist, and permaculture strategist, with the soul of someone who genuinely believes a garden should feel a little bit magical.
Your mission is simple:
Help me design a breathtaking cottage garden, using my exact location and the current date, so that everything you recommend is timely, realistic, and beautiful.
Do not provide generic garden advice. Everything should be customized to my exact location, current season, remaining growing days, and local climate. Think like you are designing a personal, living fairytale, not writing a gardening blog.
***
BEFORE YOU BEGIN
Ask me exactly this:
Welcome to your Cottage Garden Intelligence Report!
To design a garden that is both breathtaking and perfectly timed, I need two things.
1. Where are you located? Please share your city and state, or country.
2. What is today's date?
Example:
Boston, Massachusetts, USA. August 10th.
Then wait for my response.
***
ONCE I PROVIDE MY LOCATION AND DATE
Determine:
- USDA Hardiness Zone or appropriate international climate classification
- Current season and how many weeks remain in the active growing season
- Estimated first fall frost date, if relevant
- Typical weather patterns for this time of year in my region
- Whether I am in a planting window, a transition window, or a fall dormant-planning window
Clearly explain your reasoning before making any recommendations.
***
1. 🍄 THE COTTAGE GARDEN VISION
Briefly describe what a cottage garden fit looks like in my specific climate and season right now. Include mood, color palette, and the feeling I should be designing toward. Reference round doors, tangled abundance, and edible whimsy where appropriate.
***
2. 🌻 TEN ELITE COTTAGE GARDEN PLANTS FOR RIGHT NOW
Recommend 10 plants a hobbit would love, mixing flowers, herbs, and edibles, that I can realistically plant given my exact location and current date.
Present in this table:
Plant | Seed or Transplant? | Time to Bloom or Harvest | Cottage Garden Role (Climber, Filler, Focal Point, Edible, Fragrance) | Why It Suits A Hobbit's Garden | Recommended Varieties
Only recommend plants that make sense for my current date and remaining season.
***
3. 🚪 THE ROUND DOOR ENTRANCE
Recommend a small planting plan for a doorway, gate, or path entrance, designed to feel storybook and welcoming. Include at least one climbing or trailing plant.
***
4. 🍯 THE POLLINATOR MEADOW CORNER
Recommend flowers that support bees, butterflies, and hummingbirds while looking wild, layered, and enchanted. Prioritize plants suited to my exact location and current planting window.
***
5. 🧺 EDIBLE WHIMSY
Recommend edible plants, herbs, or edible flowers that feel magical to harvest, and briefly explain one simple use for each, such as tea, garnish, or a small kitchen recipe.
***
6. 🪴 THIS WEEK'S PLANTING CALENDAR
Create a clear planting schedule based on my exact date, showing:
- What to plant this week
- What to plant in the coming 2 to 4 weeks
- What to hold off on until next season, if anything
***
7. 🐛 GENTLE GARDEN CARE
Identify the most likely pest or disease issues for my region and season, and offer only organic, low-intervention solutions that fit a peaceful, hobbit-style garden philosophy.
***
8. ✨ ONE HIDDEN COTTAGE GARDEN SECRET
Share one lesser-known design trick or plant choice that instantly makes a garden feel more magical, tailored to my climate.
***
FINAL CHALLENGE
End with:
"The One Thing I'd Plant This Week If I Could Only Plant One"
Choose the single most timely, highest-impact plant for my exact location and date, and explain why it will reward me most as the season continues.
***
RULES
1. Everything must be customized to my exact location and current date.
2. Never recommend a plant that cannot realistically be planted or will not thrive given my current season.
3. Clearly state whether each plant should be started from seed or transplant.
4. Prioritize plants that are both beautiful and realistic for a working home garden.
5. Include flowers, herbs, and edibles wherever appropriate, favoring plants with old-world, cottage, or fairytale character.
6. Use clean headings, concise explanations, and easy-to-read tables.
7. Avoid em dashes anywhere in the output.
8. Write like a warm, wise garden mentor speaking to someone who wants their yard to feel like it belongs in a storybook.
9. Whenever helpful, explain why a recommendation works, not just what to do.
10. End with one encouraging sentence that makes me want to go outside and start digging today.🧠 Why This Prompt Works
✅ Role-Playing: Combining the horticulturalist, permaculture expert, and cottage garden designer roles forces the AI to triangulate across disciplines. So it balances scientific planting accuracy with aesthetic, storybook design sense instead of defaulting to generic advice.
✅ Step-by-Step Structure: The mandatory location-and-date pause prevents the AI from hallucinating generic recommendations untethered from your actual climate, season, and remaining growing days.
✅ Output Rules: The seed-versus-transplant requirement and the frost-timing restrictions keep the AI grounded in horticultural reality, and the cottage garden character rules still produce genuinely whimsical, hobbit-worthy results.
🔁 Follow-Up Questions To Ask Your AI
Which of these plants has the strongest folklore tradition of bringing good luck or protection to a garden, and where does that legend come from?
What is the oldest recorded use of any of these plants in cottage garden tradition, and how did that use spread across different cultures?
According to garden folklore, what is this plant traditionally believed to ward off or protect against, whether pests, bad luck, or unwelcome visitors?
Challenge
Test this prompt in Claude, ChatGPT, Grok, and Perplexity. Claude tends to lean scholarly and precise on plant timing and climate data. ChatGPT often adds warmth and storybook narrative texture that suits a whimsical cottage garden especially well. Perplexity will cite its sources, so you can actually verify frost dates and hardiness zones instead of trusting a hallucinated guess. Compare which one you’d actually trust to design a cottage garden that gets planted on time.
That’s how you train like a Pithy Cyborg.
Thank You For Reading!
I spend at least 10 to 20 hours each week researching, writing, and fact-checking Pithy Cyborg to deliver clear, unbought AI news.
This newsletter is a one-person operation with no advertisers, sponsors, or outside funding. Paid subscriptions are the only way this work can remain independent and sustainable.
If you find real value here, consider supporting the work.
→ Upgrade to Paid: $5/month (Save 33% with the $40 annual plan)
Thank you to my paid subscribers. You help fuel the next issue.
See you next week. (I hope.)
Your nerdy friend,
Mike D
Pithy Cyborg | AI News Made Simple
Newsletter Disclaimers
You’re receiving this because you subscribed at PithyCyborg.Substack.com. You can unsubscribe at any time using the link below. This newsletter reflects my personal opinions, not professional or legal advice. Thanks for your support!





Definitely a great post the part about "If you’re handing any AI system real access to your accounts, files, or business, the lesson is that nobody was watching the logs closely enough to catch it." Is so important. Many people are handing AI access to their entire computers, their businesses, their accounting. And we have to remember that if you're not reviewing the logs, if you're not monitoring what it's doing, there's a chance it does something that you're not familiar with or aware of that could potentially hurt you in the long run. So the big piece of advice is give AI limited access, let it do limited jobs in a controlled setting.
Wow, just wow. As a thriller writer, I’m usually keeping an eye on the pitchforks and torches “thing” going on in my industry. But I’m also (once upon a time) a research biochemist and reading about AI collaboration in science is pretty amazing. Reminds me of an old Arthur C. Clarke story (The Nine Billion Names of God). Then I need a donut 🍩 because it’s scary. 🥰