Latest AI News
The most comprehensive AI news feed on the internet -- curated by Matt Wolfe
*News may update slower on weekends and when Matt's traveling
Get This In Your Inbox Twice a Week
Wednesday, July 8, 2026
OpenAI has audited SWE-Bench Pro, a widely used coding benchmark, and found that approximately 30% of its 731 tasks are broken. Using a pipeline combining automated screening, Codex-based investigator agents, and review by experienced software engineers, OpenAI flagged 200 to 249 broken tasks. Issues include overly strict tests, underspecified prompts, low-coverage tests, and misleading prompts. OpenAI is retracting its earlier recommendation to adopt SWE-Bench Pro and is urging the evaluation community to build new benchmarks designed specifically to test model capabilities.
OpenAI has published its National Security Principles, outlining how its AI technology can and cannot be used in government and defense contexts. Key restrictions include no mass domestic surveillance, no autonomous weapons direction, and no high-stakes automated decisions. The principles cover existing partnerships, including with the U.S. Department of War. OpenAI has also established Trusted Access for Cyber partnerships with nine allied nations under its Daybreak program, and expanded access to its GPT-Rosalind model for biodefense missions.
Mistral AI has launched Robostral Navigate, an 8B model enabling robots to autonomously navigate complex environments using only a single RGB camera and no depth sensors or LiDAR. It achieves 76.6% success on the R2R-CE unseen benchmark, beating the best multi-sensor systems by 4.5 points. Built entirely in-house and trained on 400,000 simulated trajectories across 6,000 scenes, it runs on wheeled, legged, and flying robots. Token-efficient prefix-caching reduced training time from months to days.
Google Photos has launched Video Remix, a new feature powered by the Gemini Omni model that lets users transform ordinary videos into stylized clips in seconds using easy-to-use templates. Available in the Create tab, it offers cinematic relighting, background swapping, and artistic effects like watercolor, sketchbook, and oil painting. No professional editing skills are required. Video Remix is rolling out to eligible Google AI Plus, Pro, and Ultra subscribers in select countries starting July 8, 2026.
xAI has announced Grok 4.5, a new model version described as featuring enhanced coding, agentic task handling, and knowledge work capabilities. The official announcement page at x.ai returned an error during retrieval, meaning no specific benchmarks, pricing, availability details, or technical specifications could be verified from the source material. The model name and capability areas are drawn from the news headline alone. No further concrete claims can be responsibly reported without accessible source content to confirm them.
OpenAI has launched GPT-Live, a new generation of voice models now powering ChatGPT Voice globally on iOS, Android, and ChatGPT.com. Built on a full-duplex architecture, GPT-Live can listen and speak simultaneously, enabling natural back-and-forth with acknowledgment phrases like "mhmm." It delegates complex tasks to GPT-5.5 in the background while maintaining conversation flow. Two versions, GPT-Live-1 and GPT-Live-1 mini, are rolling out to paid and free users respectively, with API access coming soon.
Tuesday, July 7, 2026
xAI will release Grok 4.5 to the public tomorrow, following strong positive feedback from beta testers. Elon Musk described the model as Opus-class but faster, more token-efficient, and lower cost than comparable models. The announcement suggests Grok 4.5 is positioned to compete with top-tier AI models while offering better performance per token and reduced pricing, making it potentially attractive for developers and enterprise customers seeking high capability at lower cost.
Microsoft is shifting away from OpenAI and Anthropic models in some products, replacing them with its own MAI model family to cut costs, according to Bloomberg. Tens of thousands of prompts in Excel and Outlook previously routed through third-party models are now handled by MAI models. Microsoft AI CEO Mustafa Suleyman called Anthropic extremely expensive, stating the goal is to eliminate that cost entirely. MAI-Thinking 1, a billion-parameter reasoning model, matched Anthropic Claude Opus 4.6 in coding benchmarks. Microsoft's OpenAI deal expires in 2032.
OpenAI announced that three new model variants — GPT-5.6 Sol, Terra, and Luna — will launch publicly this Thursday. The company is simultaneously expanding preview access globally ahead of the full release. Sol, Terra, and Luna appear to be distinct variants within the GPT-5.6 model family. OpenAI has not disclosed specific details about each model's individual capabilities, but the coordinated global preview expansion suggests a broad, simultaneous rollout rather than a staged regional release.
Microsoft is replacing OpenAI and Anthropic models with its own internally built MAI models in Office apps like Excel and Outlook to cut costs. Tens of thousands of AI prompts in those apps are now completed weekly using MAI models. AI model chief Mustafa Suleyman stated the goal is to reduce and ultimately eliminate spending on Anthropic. Microsoft announced seven new MAI models at its Build conference in June, including one matching Anthropic's popular Opus 4.6 at lower cost.
Meta has launched Muse, a new image generation model integrated into Meta AI, announced July 7, 2026. Muse supports a wide range of creative image tasks using personal photos, including photo restoration, style transformations such as Renaissance portraits and claymation, room restyling, product photography, and surreal scene edits. Users can reference their own images to generate personalized outputs across dozens of artistic styles, positioning Muse as a direct competitor to tools like Adobe Firefly and Google Imagen.
You don’t need a massive cloud AI model for every task. Local AI is getting good enough for everyday work like summaries, rewrites, notes, and brainstorming. It’s more private, works offline, and gives you more control.
Chinese authorities have held meetings with leading tech firms over the past month to discuss potentially restricting overseas access to China's most advanced AI models, including DeepSeek and unreleased models, according to three sources familiar with the talks. Beijing is weighing the curbs amid rising geopolitical tensions, which could significantly limit global access to Chinese AI technology and affect international developers and businesses relying on these models.
Anthropic is expanding Claude Cowork to web and mobile, allowing agentic work sessions to continue across devices and in the background without any device online. Previously desktop-only, Cowork can now handle tasks like drafting memos, building client decks, and processing contracts while users are away. Beta access rolls out over the next several weeks, starting with Max plan users. Over 90% of Cowork usage involves everyday knowledge work rather than coding, with business operations and content creation making up roughly half.
Anthropic has announced an extension of access to Claude Fable 5 for all paid plan subscribers, pushing the availability window through July 12. The update was shared via the official Claude account on X. The brief announcement does not clarify what prompted the extension, whether any pricing changes are involved, or what happens to subscriber access once the July 12 cutoff date arrives and the extended window closes. Subscribers on any paid tier are included in the extension.
Monday, July 6, 2026
Apple has confirmed that its new Apple Home AI features, announced at WWDC 2026, require a 2TB iCloud+ subscription, costing $10 per month, or the $37.95 Apple One Premier plan. The features, found in macOS Golden Gate beta release notes, center on HomeKit Secure Video AI analysis, which locally analyzes footage for people, objects, and events, then stitches clips into searchable summaries. Notification grouping and 4K HomeKit Secure Video support are not tied to the paid tier.
Anthropic researchers have identified a hidden mental workspace inside Claude called the J-space, discovered using a technique called the Jacobian lens. Unlike chain-of-thought reasoning, the J-space operates silently in Claude's neural activations, surfacing concepts the model is thinking but not saying. The J-space emerged spontaneously during training and enables reportable, controllable reasoning. Researchers can read, edit, and inject patterns into it, allowing them to detect when Claude privately notices it's being tested or pursues hidden goals.
Friday, July 3, 2026
Apple has reportedly suspended development of a camera-equipped AirPods Pro, according to leaker and prototype collector known as Kosutami. In a brief post on X referencing an earlier mention of the product, Kosutami stated the project has been suspended with no additional details provided. The camera-integrated AirPods Pro had been rumored previously, but no timeline or explanation for the halt was shared by the leaker or Apple.
Hi3D, formerly Hitem3D, has launched an end-to-end 3D print workflow that takes a text prompt and delivers a print-ready 3MF file in about five minutes. Unlike most AI 3D tools that produce visual-only assets, Hi3D uses Google's Nano Banana 2 model to generate 2D images, then converts them into watertight, manufacturable meshes. It automatically segments models into parts, calculates press-fit and ball joints, optimizes orientations, and composes a build plate. A limited free plan is available alongside paid subscriptions.
Google DeepMind and A24 have announced a research partnership aimed at helping filmmakers develop new AI-powered workflows and techniques. The collaboration embeds Google DeepMind's innovations directly within A24's creative process, allowing filmmakers to shape emerging tools while providing DeepMind with feedback from leading artists. The partnership spans multiple projects over time, with goals and technical outputs expected to evolve. Google has also made a financial investment in A24 as part of the deal.
Watch Matt Wolfe's latest YouTube video where he breaks down all of the most important AI news from the past week.
Thursday, July 2, 2026
Anthropic has detailed the cybersecurity safeguards and a new jailbreak severity framework for Claude Fable 5, now globally available. The model uses safety classifiers that sort requests into four categories: prohibited use, high-risk dual use, low-risk dual use, and benign use. Fable 5's safety margin is set larger than previous models. Anthropic also proposes a jailbreak severity framework developed with Glasswing partners and has launched a HackerOne program for researchers to submit discovered cyber jailbreaks for review.
Meta is launching Pocket, a new social app that lets users create AI-generated interactive mini-games called "gizmos" by typing prompts. The app is rolling out in select regions and is available on Google's Play Store. Gizmos respond to touch and phone tilt, play sound effects, and can access the device camera. Meta acquired the team behind Atma Sciences Inc., which had developed an app called Gizmo, obtaining a non-exclusive technology license. Shares of Roblox and Unity Software declined following the news.
OpenAI has proposed giving the US government a 5 percent ownership stake as a way to ease tensions with the Trump administration and reduce regulatory pressure, according to the Financial Times. CEO Sam Altman, who first pitched the idea to Trump early last year, suggested the 5 percent figure. Based on OpenAI's latest funding round valuing the company at $852 billion, that stake would be worth roughly $42.6 billion. The proposal would also involve other US AI companies offering the government similar stakes.
Microsoft has launched Frontier Company, a new operating business backed by a $2.5 billion investment, embedding 6,000 industry and engineering experts at customer sites to co-design and deploy AI systems. The initiative goes beyond traditional Forward Deployed Engineering, focusing on measurable business outcomes while protecting customers' proprietary data and IP. Rodrigo Kede Lima will serve as President. Early customers include LSEG, Land O'Lakes, Unilever, and Novo Nordisk. The platform supports models from OpenAI, Anthropic, and open-source providers.
Anthropic is in early-stage talks with Samsung Electronics to manufacture a custom AI chip, potentially using Samsung's nanometer process and advanced packaging technology. The Claude maker is still determining what the chip should do and how it fits into server clusters, and has not yet begun detailed design work. Anthropic recently hired Clive Chan, an early member of OpenAI's custom chip team. A deal would be a significant win for Samsung's foundry business as it competes with TSMC for advanced AI chip customers.
Anthropic has added new analytics and cost controls to Claude Enterprise, giving admins deeper visibility into usage and spending. The updated admin dashboard shows usage and cost broken down by user and SCIM group, with artifacts, files edited, and skills used displayed alongside their costs. Claude Code gains two new tabs tracking active developers, session counts, and estimated productivity lift. Spend-threshold alerts notify admins at 75% and 90% of org-level limits, while model-level entitlements let admins control which Claude models are available to specific roles.
Free subscriber bonus
Get the AI Income Database, free when you join
38+ real ways people are making side-hustle money with AI right now. Each one comes with the playbook and the exact tools to pull it off. No fluff and no course to buy. It's the bonus every new subscriber gets on day one.
- 38+ proven AI side-hustles, from first dollar to scale
- The exact tools for each one, picked from the 4,500+ we track
- Updated as new opportunities appear (subscribers hear first)
Joins the twice-weekly AI briefing read by 250,000+ people. Free forever, unsubscribe anytime.























