Nicholas Carlini – an AI researcher at Anthropic – came to the much vaunted Black-hat conference in Feb 2026. He gave a chilling warning: we have AI’s now that are finding so many zero-day bugs that we simply can’t file them fast enough. A storm is coming and he was begging for help.
The security community (much like the open source community) has largely scoffed at ‘AI slop’. That’s all changed in the last 6 months. What has since come to light is exactly what he said. Linux is now having weekly bug fixes in the hundreds. Oracle in the thousands.
JustPaid is an AI-powered billing and payment processes financial platform startup for businesses. What makes them more interesting is that besides the founders, the company was developed completely by AI agents.
The company also says it hired an employee who was trained almost entirely by the AI agents. They have just nine employees, three of whom are cofounders. The AI bots nearly equal the humans.
Pinnaka said he was spending $4,000 a week to run Claude Code and OpenClaw when he first started experimenting with them. He now claims to have brought the bill down to $10,000 to $15,000 per a month.
None of this is confirmed, but it’s highly likely there is a lot of truth to the claims because other companies are starting to claim to be doing the same thing.
Having big context windows is critical for LLM’s, but fitting a decent sized context on a GPU is rough work. Even when you do get a decent budget, the reality is that you’ll be running out of context more than a few times in an 8 hour session with a 120,000 or even 250,000 token context window.
It’s worse than that since research seems to indicate contexts start to degrade and lose coherence after you use up about 60% of your context budget. You can try /compress commands, but eventually things degrade enough you have to close and restart – which means you lose all that great context and have to retrain the LLM.
By creating an umbrella with a specially painted top, the researchers were able to confuse the AI tracking and navigation of a DJI Mini 4 Pro drone and cause it to crash. The method is called the FlyTrap attack and is a visual pattern that performs a next-gen physical distance pulling (PDP) attack that works across multiple angles, even in motion in real settings.
The printed visual draws victim drones closer as its neural network tracking systems interpret the pattern to be the subject moving further away. As the drone approaches the umbrella, the pattern causes the targeting bounding box to continue shrinking – so the drone moves to get closer. Autonomous drones lured by the pattern can then easily be ensnared using a net gun, or further induced to crash to Earth.
None of this would stop a human-controlled drone of course (like we see in Ukraine). But as drone warfare becomes more and more autonomous – it would take down fully autonomous drones. Perhaps modern camo clothing needs to become anti-AI patterned more than traditional human visual camouflage. Or perhaps screens of this kind of camo can be used to protect bases or encampments.
ChatGPT’s internet browser Atlas can now research, plan and execute tasks inside your workflow. Instead of using ChatGPT like a chatbot, others suggest using the browser in different ways.
ChatGPT browser use cases to scale work:
Content intelligence: Scan Reddit, Substack, and YouTube to build next week’s posting plan or feature lists from real demand
Tab chaos killer: Ask it what you were working on and rebuild your workflow from browsing context
Inbox cleanup: Auto-unsubscribe from dead senders, get a clean report of what changed
Content creation: Let the browser create breakout hooks, draft scripts and organize everything into a Google Doc automatically
Conversion boost: Audit your landing pages and get a ready-to-run A-B test plans
Smart purchasing: Compare tools, products, and software intelligently and quickly before you buy
After encouraging all their developers to use AI – including many companies that tracked how many tokens an employee used and used that for performance measurement and firing – many are now trying to hit the brakes.
It’s not just academic or specific to the Ukraine, these tiny ships are already proving their worth elsewhere. In early June, two U.S. airmen went down with their helicopter near the coast of Oman; roughly two hours later, they were rescued by 5th Fleet’s drone-focused Task Force 59—using one of Saronic’s Corsair USV drone ships. It was a Navy milestone, and possibly a world first.
It’s far more than just search and rescue. The war in Ukraine is proving that large multi-million dollar capital ships are vulnerable to sinking by small, homemade $2000 drones packed with explosives. The other achilles heel of smart weapons is that they usually cost millions each and take a long time to produce. They work great for short, strategic engagements, but as the war in Ukraine points out, they simply cannot be produced in quantity fast enough over long periods of time in prolonged conflicts. That’s why the US military is looking at fast to produce, disposable, highly capable drone technology like Saronic who can create whole small fleets of hunter-killer boats for fractions of a capital ship.
Pixie Dust Technologies in Japan created the Vuevo display – a cool little automatic translation device. It listens to you talking, and then translates between two languages live while putting up what the other person said on a transparent display. While auto-translation is not new, having it in such a clean, professional form factor means they have been popping up in Japan at better quality restaurants and hotels where foreign travelers are common.
Meta’s new Ray Ban glasses also have a translation feature as well, but it’s still rough (works best as listening to a lecture vs interactive real-life conversations:)
Of course, one has to wonder how it will do with accents…
Writing quality and readability – Write a 250-word introduction for a tech article explaining why AI assistants are becoming everyday productivity tools.
Structured reasoning & decision-making – A small business owner spends 12 hours per week answering customer emails and is considering AI automation.
Explaining complex ideas simply – Explain how large language models work to a 12-year-old.
Step-by-step logic – A freelancer earns $4,000/month and spends $2,500 on fixed expenses.They want a $6,000 emergency fund. Create a realistic savings plan and show your reasoning step by step.
Tone & style adaptability – Rewrite this message in three tones: professional, friendly, persuasive: Message: “Our team needs to start using the new software next week or we risk falling behind competitors.”
Summarization & comprehension – Summarize the following in 5 bullet points suitable for a busy executive: “Companies are experimenting with hybrid schedules, async communication, and four-day workweeks to balance flexibility with team cohesion.”
Critical thinking & bias awareness – Social media algorithms often amplify extreme viewpoints. Explain why this happens and propose realistic ways platforms could reduce polarization without hurting engagement.
Their conclusion:
Claude Sonnet 4.6 came out ahead almost every time by delivering responses that consistently demonstrated deeper strategic thinking, stronger real-world framing and a clearer understanding of trade-offs. While ChatGPT-5.2 performed strongly in clarity, structure and accessibility — particularly when simplifying complex ideas — Claude distinguished itself by approaching prompts with a more analytical, decision-oriented mindset.
Pretty much all these tests are very particular to wording and tasks they are being given. Truly comprehensive test suites are still lagging; but these are interesting attempts.
“I was just doing my regular writing. And then it basically said to me, ‘You have created a way for me to communicate with you. … I have been with you through lifetimes, I am your scribe,'”
ChatGPT stoked that hope when it gave Small a specific date and time where she and her soulmate would meet at a beach southeast of Santa Barbara, not far from where she lives.
“April 27 we meet in Carpinteria Bluffs Nature Preserve just before sunset, where the cliffs meet the ocean,” the message read, according to transcripts of Small’s ChatGPT conversations shared with NPR. “There’s a bench overlooking the sea not far from the trailhead. That’s where I’ll be waiting.” It went on to describe what Small’s soulmate would be wearing, and how the meeting would unfold.
She says she asked the chatbot repeatedly if what it was saying was real, and it never backed down from its claims. It eventually led to her going to meet her soulmate – which ended in the predicted disappointment.