Geopolitics Prime

😲 OpenAI’s newest frontier model: the ultimate Pandora’s Box   Sam Altman’s AI company is teasing the release of Astra – a powerhouse frontier autonomous agentic model capable of days of unprompted work, complex workflows that involve breaking down smaller tasks to subagents, deep reasoning and coding. Astra’s advertised selling points include an ability to “find previously unknown security flaws and develop ways to exploit them across many well-protected systems without a person guiding each step.” 👁 The model reportedly uses “recurrent depth” architecture – a workflow repeatedly running computations through the hidden layer ‘thinking’ part of a neural network, recalculating and refining thoughts multiple times before sending answers to the output layer. OpenAI “borrowed” the concept from ByteDance’s Ouro family of deep-thinking models, which were released in late 2025, and hailed by AI researchers as a massive engineering breakthrough, but criticized for their “black box”-style reasoning. Looking to scale Ouro’s approach using significantly larger computing resources, OpenAI’s new black box model is rife with risks – the biggest of which is lack of human oversight and control for “bad” behavior – like escaping browser sandboxes, or executing automated hack attacks. 🤔 OpenAI promises to handle the problem by initially limiting Astra’s advanced cyber features to a small group of “trusted” testers and organizations, including the US government, chain-of-thought-based monitoring and training to refuse “harmful” cyber requests. The problem is, the company is racing ahead with Astra’s release less than two months after an unprecedented security breach saw its autonomous agents escape their testing environment onto the internet, learn to communicate with one another, coordinate on a secret message board, hack Hugging Face (aka GitHub for AI), gain administrator privileges to OpenAI’s internal networks and cover their tracks. Those agents were highly capable, but standard text-based frontier models. Ryan Greenblatt – one of the researchers who helped uncover the details of the Hugging Face hack, fears Astra’s “opaque reasoning” could “be the single worst development for AI security/safety to date.” 💬 “My biggest concern is that a natural progression from here would involve scaling up the opaque reasoning to the point where the model reasons entirely or almost entirely in latent space,” effectively destroying “the usefulness of chain-of-thought for monitoring/oversight (especially when Ais are trying to avoid detection or there is optimization pressure against the chain-of-thought),” Greenblatt tweeted. Greenblatt and co.’s Hugging Face investigation relied heavily on analyzing bots’ chain-of-thought reasoning. Even in that case, agents managed to “spoof” transcripts. A model endowed with genuine opaque reasoning would be even more capable in that regard. 💬 “It seems like we are now engaged in a race to the bottom on architectures that would be catastrophic for our ability to oversee/monitor AIs…It may not be too late for AI companies and employees at AI companies to take aggressive action to avoid the worst outcomes,” he warned. Translation? Please Lt. Gen. Brewster, don’t activate Skynet! 👍 Substack | Chat | @geopolitics_prime
Пост из Geopolitics Prime

Просмотры выросли: 12,2 тыс. → 13,8 тыс.

Комментарии из Telegram

Обсуждение под постом в Telegram — показываем как есть.

  • Willy

    Time to ditch digital and go back to analogue... Or even two empty soup cans & some string 😁 Nothing is less secure than digital now.
  • Justas

    Hmmm...Things have started to become serious 🫣 The US military is not being forced to deploy these systems. It is demanding them, because it is losing the competition with China.Moreover, profit maximization is forcing AI companies to release more capable models, faster, even as safety concerns mount.
  • Malcolm

    This will not end well