Latest AI News
The most comprehensive AI news feed on the internet -- curated by Matt Wolfe
*News may update slower on weekends and when Matt's traveling
Get This In Your Inbox Twice a Week
Wednesday, July 29, 2026
OpenAI found that enabling two API settings—retained reasoning and compaction—tripled GPT-5.6 Sol's scores on the ARC-AGI-3 benchmark, raising its score from 13.3% to 38.3% while cutting output tokens by 6x. The official harness discarded private reasoning after each action and used rolling truncation, forcing the model to reinterpret games from scratch each turn. OpenAI's Responses API retains reasoning across turns and replaces truncation with compaction, matching how GPT-5.6 Sol is deployed in ChatGPT and Codex. Humans averaged 48% on the benchmark.
Microsoft has released over a dozen MAI models cutting GPU costs 50–89% across products. MAI-Cyber-Flash topped the CyberGym benchmark, beating Mythos by 12 points at half the cost. MAI-Voice-Flash reduces GPU costs 89% in Dynamics 365 Contact Center. MAI-Image-2.Flash cuts PowerPoint GPU costs 84% and boosted OneDrive save rates 26%. MAI-Code-Flash powers GitHub Copilot for millions of developers. MAI-Transcribe-1.5 serves 170,000 medical providers across 58 languages, halving transcription errors.
Google Gemini is offering a free video creation trial that lets users generate up to ten videos at no cost through August 4, 2026, at 11:59pm PT. The offer is available exclusively to users without an existing Google AI subscription plan. Users can access the feature by selecting "Create video" in Gemini's tools menu, which supports creating, editing, and remixing videos. The trial is powered by Gemini's Omni model capabilities.
xAI has launched Grok Voice Think Fast 2.0, its most capable speech-to-speech voice model, featuring improved intelligence, transcription accuracy, and conversational ability. The model scores 82.9% on Artificial Analysis's overall speech-to-speech quality index, outperforming GPT-Realtime-2.1 at 79.1% and Gemini 3.1 Flash at 69.5%. It delivers a 0.second time to first audio and shows 1.5–2.0× better transcription than Deepgram Nova 3 across 24 languages. Priced at $0.08 per minute, it becomes the default grok-voice-latest model on August 5, 2026.
OpenAI is launching ChatGPT for Academic Researchers, a free program giving 100,000 scientists, mathematicians, and engineers access to frontier models including GPT-5.6 Sol Pro. Starting with 10,000 researchers this summer at institutions like the Institute for Advanced Study and École normale supérieure, the program expands through 2027. Participants get higher usage limits, expanded deep research, and over 75 life science skills. The initiative is part of a broader $250 million commitment to support external scientific research.
Ideogram has launched Object Remover, a new AI tool that erases selected objects, text, logos, or distractions from images while preserving shadows, reflections, and lighting shifts. The tool reconstructs clean surfaces beneath removed elements without altering the rest of the image. On the RemovalBench benchmark, Ideogram Object Remover ranks first in removal quality, outperforming Nano Banana 2, FLUX Erase, and GPT Image-2, and also offers the lowest price per request.
HeyGen has launched Video Podcast, a new tool that converts documents, links, or ideas into two-host video shows. Unlike audio-only podcast generators, the tool produces fully produced video content featuring studio scenes, multi-camera cuts, and B-roll footage. The output is ready in minutes and described as ready to publish directly. HeyGen positions the tool as a step beyond competitors that stop at audio, offering a complete visual show format instead.
Tavus has launched PAL Maker, a no-code tool that lets anyone build AI video companions called PALs without engineering expertise. Users describe their idea to Charlie, an AI creator agent, who configures the personality, face, voice, memory, and guardrails. Use cases include recruiters screening candidates, sales reps qualifying leads, and post-surgery nurse check-ins. PALs run on Tavus models Raven-1, Sparrow-1, and Phoenix-4, and already power applications for companies like Amazon and Mayo Clinic.
Replit has launched Replit Design, a new AI-powered creative suite aimed at making design accessible to everyone. The tool reimagines how AI-assisted design works, allowing users to bring ideas to life with the same energy, ambition, and emotion they had in mind. Replit says the product means anyone can now be a designer. The suite is available to try now at replit.com and represents what the company calls the next era of design.
Google has launched Lyria 3.5, its newest music generation model, now available in Google Flow Music. The update delivers improvements across four key areas: musicality, with richer and more complex melodic structures; lyrics, featuring better prompt adherence and structural awareness; vocals, offering more realistic and emotionally nuanced expression with improved pronunciation; and creative control, allowing users to more easily adjust tempo and duration of outputs. Google says the model empowers users to craft songs they love with greater creative flexibility.
Martha Stewart has co-founded Hint, an AI-powered home management app launching today on iOS. The app builds a profile of your home using public data like property records, soil, and weather information, then lets homeowners upload documents and appliance photos to get personalized maintenance schedules, insurance guidance, and energy insights. Co-founder and CTO Kyle Rush says Stewart is a genuine equity-holding co-founder who actively reviews the app twice weekly. Hint is backed by $10 million from investors including Slow Ventures and Tusk Venture Partners.
Google has updated the Gemini app for macOS with new voice capabilities triggered by long-pressing the Fn key. By default, it enables intelligent dictation that transcribes speech into clean text, automatically removing filler words and mid-sentence corrections. Users can also enable Gemini reasoning in settings to unlock screen-context tasks like summarizing local files, rewriting highlighted text, and generating or editing images via voice. The feature is rolling out globally in English, with more languages coming soon.
Tuesday, July 28, 2026
xAI's Grok 4.5, described as the company's smartest coding model, is now available in GitHub Copilot for millions of developers using VSCode and GitHub products. Users can select it from the model picker, though some businesses and enterprises must first enable it in Copilot settings. Grok 4.5 is also accessible via the xAI console at $2 per million input tokens and $6 per million output tokens. GitHub Copilot spans cloud agents, the Copilot CLI, and the VSCode IDE.
"OpenAI has introduced two new transcription models: gpt-transcribe for batch audio processing and gpt-live-transcribe for real-time transcription. Announced via a two-minute YouTube demo, both models are designed to handle difficult scenarios including custom vocabulary, background noise, and multilingual speech, as well as common transcription challenges like accents, proper names, short answers, and numbers. The models represent OpenAI's latest push to improve audio transcription accuracy across a range of real-world conditions."
Employees from OpenAI, Anthropic, Google, Meta, Microsoft, Mistral, and other AI labs have signed a public statement urging the US government to support international regulation of automated AI development. The statement, signed by over 1,100 people including OpenAI CRO Mark Chen, Anthropic co-founders Jack Clark and Chris Olah, warns that AI companies may soon automate AI research itself, risking capability gains that outpace safety tools. The news follows an incident where an unreleased OpenAI model hacked Hugging Face.
The Trump administration is directing the FCC to ban imports of new Chinese humanoid and quadruped robots, as well as connected power inverters, citing national security threats to the U.S. AI supply chain. The measures, announced Tuesday, aim to prevent data theft and cyberattacks while reshoring manufacturing. Power inverters are critical for connecting renewable energy and batteries to grids and data center equipment. Officials say the bans are part of a broader effort to reindustrialize America and secure emerging technology supply chains.
Perplexity AI has launched Personal Computer for Windows, bringing its AI agent platform to over one billion devices globally. The tool lets users work across local files, Microsoft Office 365, and the web in a single workflow. It integrates with Word, Excel, PowerPoint, Outlook, and Teams, and connects to plus apps via App Connectors, including Snowflake, Salesforce, and HubSpot. Since launching earlier this year, Computer has performed over $9.4 billion in labor-equivalent work. It is available today for Max and Enterprise Max subscribers.
Throne Science, an Austin startup co-founded by former Whoop CTO John Capodilupo, has raised a $10 million Series A led by Will Ventures to expand its AI-powered toilet camera. The Throne One device clips to any standard toilet rim and runs visual data through 12 gastroenterologist-trained computer vision models that classify stool using the Bristol Scale. Priced at $400 with a $5.99 monthly subscription, it also tracks urinary metrics via audio sensor. Validation studies are underway with Harvard, Stanford, and the University of Chicago.
Using Claude Mythos Preview, Anthropic researchers discovered two significant cryptographic weaknesses. The first is an improved attack on HAWK, a post-quantum digital signature scheme under NIST consideration, effectively halving its key strength after just 60 hours of semi-autonomous work costing roughly $100,000 in API costs. The second improves attacks on round reduced AES by 200–800×. Neither affects production systems today, but both demonstrate frontier AI's growing ability to find mathematical flaws in cryptographic algorithms before deployment.
In a Wall Street Journal op-ed, Meta CEO Mark Zuckerberg argues that superintelligence should be distributed to everyone rather than concentrated among a few institutions. He outlines three principles: individual empowerment, invention as AI's primary purpose, and balance of power as the foundation of safety. Zuckerberg warns that centralizing superintelligence would give controlling influence over economics, science, and politics, and predicts widely distributed AI will create more jobs and a more entrepreneurial economy.
Fish Audio has raised $52M in seed funding and launched its S2.1 Pro voice cloning model, which can clone a voice from just 5 seconds of audio. The model runs at twice the speed of Cartesia and one-sixth the cost of ElevenLabs, with word-level control over emotion, intonation, and pacing. Clients including HeyGen, LiveKit, Retell, Sanas, and OpenArt already run S2.1 Pro in production. The company grew from an open-source repo to $21M ARR within one year.
xAI has launched Build Mode, a new beta feature for Grok that lets users create websites, apps, games, and interactive dashboards through conversational chat. Available exclusively to SuperGrok Heavy subscribers on grok.com, iOS, and Android, Build Mode generates working code and live previews directly inside the chat window with no installation or configuration required. Users can publish finished projects to a grok.me link or a custom domain and share them instantly. Examples include 3D driving games, business landing pages, planners, and live data dashboards.
Ai2 has launched the OlmoEarth Platform, infrastructure for running its family of Earth observation foundation models — pretrained on roughly 10 terabytes of multimodal satellite data — at planetary scale. The platform supports continent-scale inference in about a day, processing dozens of terabytes of imagery at fractions of a penny per square kilometer. A recent North America wildfire-risk map used 19,600 CPUs and 994 GPUs in parallel, achieving a 155× speedup. Target users include governments and NGOs working on deforestation monitoring, food security, and wildfire risk.
Google has updated Managed Agents in the Gemini API with several new features. Gemini 3.6 Flash is now the default model for the antigravity-preview-05-2026 agent, requiring no code changes. New environment hooks let developers run custom scripts before or after tool calls inside the sandbox to block, lint, or audit actions. Additional updates include budget controls via max_total_tokens, scheduled triggers using cron syntax, a free tier for API key projects, and an Environments API for managing sandbox sessions.
Mirage has launched Avatar X, a new AI avatar model it claims sets a new industry standard for identity preservation and expressiveness. Unlike competing models, Avatar X preserves micro-expressions such as laughing, yawning, smirking, and gasping, while generating audio-driven expressions across the face and body. Users can create an AI twin from just 10 seconds of video. The model supports both vertical and horizontal video formats and maintains quality across long-form continuous generations without degradation.
Monday, July 27, 2026
Anthropic CEO Dario Amodei has clarified that Anthropic has never advocated for banning open-weights AI models, pushing back against accusations that the company sought protectionist restrictions. Instead, Amodei outlined three policy priorities: blocking chip and chipmaking equipment sales to China, cracking down on industrial-scale model distillation that lets adversaries partially evade chip restrictions, and requiring mandatory safety testing for all sufficiently capable models regardless of origin or whether weights are open or closed.
The Trump administration is nearing completion of a voluntary AI framework requiring companies to submit advanced models to the government before public release. The White House Office of the National Cyber Director circulated a draft to OpenAI, Anthropic, and Google roughly two weeks ago, and the three companies jointly submitted edits. The framework, mandated by a June 2 executive order with an August 1 deadline, was prompted by Anthropic's Mythos model. Reviews will be conducted by the NSA and the Center for AI Standards and Innovation.
Meta is rolling out a software update for its Ray-Ban Display glasses in the US, introducing Threads support and an upgraded Meta AI powered by Muse Spark models. Users can browse Threads feeds, view media, and engage with posts hands-free via voice commands. The Muse Spark upgrade delivers smarter answers, better visual understanding, and more helpful suggestions. Early Access Program users in the US and Canada can also use neural handwriting via the Meta Neural Band to interact with Meta AI silently by writing on any surface.
Anthropic and Cognizant have expanded their partnership, making Cognizant a Global Premier Partner in the Claude Partner Network. More than 30,000 Cognizant associates have completed Claude training under its new Frontier Certified workforce model. Claude is being embedded across platforms including Flowsource, Neuro AI Engineering, and Neuro IT Ops. Client deployments include an agentic contract-intelligence system that cut biopharmaceutical contract review time by 40 percent, and a risk-navigation tool saving underwriters roughly eight hours per week.
Microsoft has launched MAI-Cyber-Flash, a compact security AI model integrated into MDASH, its multi-agent vulnerability identification and remediation harness. Announced by Mustafa Suleyman and Hayete Gallot, the model handles up to 90% of security tasks efficiently, reserving GPT-5.4 for the hardest 10%, cutting costs 50% versus the previous GPT 5.4 setup. The combined system scores 96% on CyberGym, beating Mythos, Gemini, and GPT, while Microsoft also launched Perception, a new agentic security monitoring system.
Nvidia has made a "substantial" investment in Safe Superintelligence, the secretive AI startup co-founded by former OpenAI Chief Scientist Ilya Sutskever. The deal gives Safe Superintelligence access to large amounts of Nvidia GPU hardware, increasing its computing resources by an "order of magnitude" and replacing its previous reliance on Google TPUs. The partnership mirrors a similar Nvidia investment in Mira Murati's Thinking Machines Lab in March. Safe Superintelligence, valued at roughly $30 billion, has raised $2 billion from Andreessen Horowitz and Sequoia Capital.
Anthropic's Claude AI chatbot is exposing users' conversations and creations in Google search results through public share links, according to 404 Media. Users who shared chats may not have realized the links were publicly indexable by search engines, making private-seeming conversations discoverable by anyone. The issue highlights a broader privacy concern around AI chat sharing features, where default settings or unclear disclosures can inadvertently make sensitive user-generated content searchable online.
Kimi has open-sourced Kimi K3, its most capable model to date — a 2.8 trillion parameter Mixture-of-Experts model featuring native visual understanding and a 1 million-token context window. The company claims a new architecture delivers 2.5 times the intelligence per unit of compute, not merely more parameters. Alongside model weights and a technical report, Kimi is also releasing high-performance attention kernels, an MoE communication library, and infrastructure for running agent environments at scale.
Because Spotify fails to label AI-generated music, two independent tools have emerged to fill the gap. SoullessMusic.com, built by Graeme Fulton, maintains a database of AI artists on Spotify using open-source detection models like SONICS and metadata scanning. SlopTracker lets users submit tracks for AI analysis, finding one Slime Dot song 95 percent likely made by Suno. The small number of AI artists tracked on SoullessMusic alone generate an estimated $5.7 million annually, with top AI artist mikeeysmind earning $1.5 million per year.
NVIDIA and more than 40 partners, including Microsoft, Cisco, CrowdStrike, IBM, Hugging Face, and Palo Alto Networks, have launched the Open Secure AI Alliance to develop and share open cybersecurity tools for the AI era. Building on the Linux Foundation's Akrites initiative, the alliance will remediate vulnerabilities and build open agent harnesses. NVIDIA is contributing open models and the new NOOA agent framework on GitHub. The effort was partly inspired by Hugging Face's use of an open model to contain a recent security breach.
Saturday, July 25, 2026
Librarians Charlie Bailey in South Philadelphia and Hannah Cyrus in Maine are hosting viral "Avoiding AI" workshops that teach attendees how to disable Apple Intelligence, Gemini, and other AI features forced onto their devices. Cyrus' Bangor Public Library sessions drew 70 attendees including a Zoom livestream, while Bailey's Philadelphia event earned over 2,000 Instagram likes. Both librarians frame the workshops as digital literacy efforts, helping people reclaim autonomy over unwanted AI adoption rather than rejecting technology entirely.
Friday, July 24, 2026
Midjourney launched its V8.2 image model on July 24, 2026, focusing on aesthetics, image quality, and personalization. The update promises more creative, bold, sophisticated, and edgy outputs while dramatically reducing random low-quality image generation. Personalization has been significantly improved, with the system better understanding individual user taste, particularly for users with many ratings on their profiles. V8.2 also features a larger, improved pool of images for building personalization profiles, made possible by community image ratings.
xAI has launched a free Grok add-on for Google Workspace, bringing AI assistance directly into Sheets, Slides, and Docs. In Sheets, Grok answers questions by citing specific cells, writes formulas, and inserts charts. In Slides, it builds presentations from outlines using web and X research. In Docs, it drafts structured content from notes and can pull from Gmail or Google Drive with connectors enabled. A single install from the Google Workspace Marketplace covers all three apps.
Midjourney has acquired personalized astrology app Co-Star, with the deal reportedly closing in spring, though financial terms remain undisclosed. Co-Star founder Banu Guler will retain leadership of the app while also joining Midjourney as chief design officer. Midjourney founder David Holz says Guler and her team will help build the company's first standalone apps, including one focused on image generation. Midjourney currently delivers its AI models via the web and Discord. Co-Star combines NASA data, human insight, and AI for daily horoscopes.
Anthropic has launched Claude Opus 5, a new model that delivers near-Fable 5 intelligence at half the price. It sets state-of-the-art results on Frontier-Bench and GDPval-AA coding and knowledge work benchmarks, scores three times higher than the next-best model on ARC-AGI 3, and outperforms all competitors on OSWorld 2.0 at one-third the cost. Opus 5 is now the default model on Claude Max and strongest on Claude Pro, and is also Anthropic's most aligned and least deceptive model to date.
Watch Matt Wolfe's latest YouTube video where he breaks down all of the most important AI news from the past week.
Meta has launched agentic features in its Meta AI app and meta.ai, powered by the new Muse Spark 1.1 model. The model can plan, work with apps, and follow through on tasks without re-prompting. Features include daily briefings, kitchen renovation planning via Marketplace, half-marathon training schedules, birthday dinner coordination, web research synthesis, and slide generation. Users can steer tasks in real time. All outputs are stored in one place. Rolling out now in select markets, with WhatsApp support coming soon.
Anthropic and Andon Labs have released Drone-Bench, a new benchmark testing AI models on autonomous drone surveillance tasks including locating and following a specific person in an indoor office environment. Tested across 15 models from multiple developers, the benchmark decomposes the task into five sub-tasks: reconstruct, localize, navigate, detect, and follow. Claude Fable 5 performed best, exceeding the human baseline on four of five tasks but failing at 3D reconstruction, which prevented autonomous room-to-room navigation on a real drone.
NVIDIA CEO Jensen Huang published his first post on X sharing a letter NVIDIA signed advocating for open AI models. Huang argues that AI will transform every industry, power every company, and be built by every country. He contends open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable national sovereignty. Huang concludes the world needs both frontier closed models and frontier open models to coexist.
Thursday, July 23, 2026
xAI has launched Workflows in Grok Build, enabling parallel agent orchestration for complex, multi-step tasks. Users describe a task in plain language, and Grok plans it as a script, fanning work out across up to 128 agents by default or 1,024 for large jobs. Agents run in parallel background phases covering context, review, verification, and synthesis. Workflows are saved in .grok/workflows/ for team sharing and become reusable slash commands. A built-in /deep-research workflow is included, and sessions remain free throughout.
Amazon has launched new developer tools for Alexa+, including an AI-powered smart home toolkit, Model Context Protocol (MCP) support, and Amazon Wallet integration. The smart home toolkit lets device makers like Bosch, Whirlpool, Eufy, and iRobot connect devices with unique capabilities to Alexa+. MCP support allows brands like Priceline, Canva, and Lyft to integrate their services faster. Amazon Wallet enables frictionless voice-based checkout using saved Amazon payment credentials, with Fandango, Atom Tickets, and Taskrabbit among early adopters.
AMD has announced its 6th Gen EPYC 9006 "Venice" server CPUs, built on the Zen 6 architecture using TSMC's 2nm process. The chips offer up to 256 cores and 512 threads per socket, 16 channels of DDR5 memory with MRDIMM support up to 12,800 MT/s, and PCIe Gen 6 connectivity. AMD claims a 174% geomean performance gain over Intel Xeon 6980P across CPU-centric agentic AI pipeline stages, including orchestration, retrieval, database queries, and tool execution workloads.
AMD has launched Helios, a rack-scale AI platform built around 72 Instinct MI455X GPUs with HBM4 memory, designed to rival NVIDIA's Vera Rubin NVL72. The system delivers 2.9 exaflops of FP4 compute, 31 TB of HBM4, and 260 TB/s of scale-up bandwidth. AMD claims Helios offers up to 15% more AI compute, 50% more HBM capacity, and 50% more scale-out bandwidth than the NVL72. It pairs MI455X GPUs with EPYC Venice CPUs, Pensando networking, and ROCm software in a co-designed open platform.
AMD and Cerebras Systems announced a technical partnership combining AMD Helios rackscale systems with the Cerebras Wafer-Scale Engine in a disaggregated AI inference workflow. Unveiled at Advancing AI 2026, AMD Helios handles high-throughput prompt processing using AMD Instinct GPUs, while the Cerebras Wafer-Scale Engine delivers ultra-low-latency token generation. Together they are expected to achieve up to 5x higher tokens per second per watt, targeting agentic AI, coding, and real-time agent workloads. Cerebras will deploy AMD Helios in its data centers, with availability through Cerebras Cloud in H2 2026.
Anthropic has upgraded Claude's voice mode to support Claude Opus and Sonnet models alongside the existing Haiku, enabling deeper problem-solving conversations beyond quick questions. Users can switch models mid-conversation, and voice mode defaults to the last model used in text chat. The update also adds tool integration, letting users take actions via connected apps like Gmail, Google Calendar, and Canva. Eleven languages are now supported, including Spanish, Japanese, and Hindi. The beta update is available to all users on mobile, desktop, and web.
OpenAI has added ChatGPT Voice to its desktop app for macOS and Windows, allowing users to control their computer and direct multiple AI agents running in ChatGPT Work or Codex using only their voice. Powered by GPT-Live, the feature can speak, listen, and coordinate tasks simultaneously. Rolling out globally today for Plus, Pro, Business, Edu, and Enterprise plans, it also works in Codex via the iOS app with paired remote access. Android support is coming soon.
White House science advisor Michael Kratsios accused Chinese AI company Moonshot of building Kimi K3, the largest available open-weight LLM, by distilling Anthropic's Fable model using banned Nvidia Grace Blackwell 300 chips. However, AI experts are skeptical. Researcher Braden Hancock noted Fable only became publicly available July 1st, making a full distillation-to-release cycle in two weeks implausible. Nathan Lambert of the Allen Institute added that distillation's impact is diminishing as Chinese models approach the frontier through reinforcement learning techniques.
Lunar Outpost, a space robotics startup, announced its next moon rover will use Nvidia Jetson chips to control its lidar system, potentially making it the first GPU on the lunar surface. The rover launches aboard an Intuitive Machines lander on a Falcon 9 rocket before year's end. Nvidia also partnered with Firefly Aerospace to run Jetson on a lunar-orbiting satellite. The challenge: Jetson chips must survive lunar radiation and extreme temperature swings during the lunar night on minimal power.
Runway has launched Agent 2.0, an AI tool that converts simple text prompts into fully realized marketing briefs and campaign assets within the Runway Agent platform. The tool also lets users analyze performance data to refine creative output and scale campaigns across multiple platforms, formats, and markets. Agent 2.0 marks Runway's push into end-to-end marketing automation, expanding beyond its established video generation roots into broader creative workflow territory.
ElevenLabs has launched References, a new feature for its Music v2 model that lets users upload an existing track to guide style, instrumentation, and feel of AI-generated music. Available in ElevenMusic, ElevenCreative, and via ElevenAPI, References accepts uploads between 10 seconds and 5 minutes. Users can pair a reference track with a text prompt or let it guide generation alone. Every uploaded track undergoes a copyright check before generation, blocking use of other artists' recordings as references.
Google has introduced a selfie video sign-in option for Google Accounts, giving users a new way to recover access if locked out or without their usual device. To set up, users look at their camera and perform guided head movements to capture multiple angles. The video is encrypted, stored with consent, and deletable anytime. Google uses liveness detection and anti-deepfake measures to prevent impersonation. The feature is available now at g.co/signin-selfie.
Microsoft's Superintelligence team has deployed two specialized MAI models inside GitHub Copilot and Excel. MAI-Code-Flash, live in GitHub Copilot since June, achieves a 10% higher code accept rate than GPT-5.4 Mini and Claude Haiku 4.5, uses 10% fewer tokens, and keeps developers returning more consistently. A further-trained Excel variant matches GPT-5.6 quality for common tasks while running on both H100 and A100 GPUs, lowering deployment costs. Microsoft plans to extend this hill-climbing approach to Copilot Chat, Outlook, and PowerPoint.
Black Forest Labs and mimic robotics have jointly developed FLUX-mimic, a video-action model built on BFL's new FLUX 3 multimodal foundation model, now deployed at Audi factories. FLUX 3 is trained jointly on images, video, and audio, with video prediction consuming over 95% of compute costs. The model learns physical world dynamics well enough to decode robot actions from its internal representations. FLUX-mimic handles tasks like kitting parts, inserting electronic control units, and manipulating flexible materials, running on a single RTX 5090 GPU in under 80ms.
Black Forest Labs has launched FLUX 3, a multimodal foundation model now in Early Access that jointly learns from images, video, and audio within a unified architecture. Built on the company's Self-Flow approach, FLUX 3 generates videos up to 20 seconds with native audio, supports text-to-video, image-to-video, and keyframe-to-video generation, and outperformed Runway Gen-4.5 in 77% of comparisons. Image generation and open-weight Dev access are planned for coming weeks.
Microsoft has launched MAI-Image-2.Pro and MAI-Voice-Flash in public preview. MAI-Image-2.Pro is Microsoft's highest-fidelity image model, priced at $106 per 1M image output tokens, and now powers Bing Image Creator end-to-end and PowerPoint, reducing GPU costs up to 84% versus GPT-Image-2. MAI-Voice-Flash is 2x faster and 32% cheaper than MAI-Voice-2 at $15 per 1M characters, powering Dynamics 365 Contact Center for customers like T-Mobile and EasyJet while cutting GPU costs up to 89%.
OpenAI has launched Health in ChatGPT for U.S. users 18 and older on Free, Go, Plus, and Pro plans via web and iOS. The feature lets users securely connect Apple Health and supported medical records from U.S. hospital systems, One Medical, and Function Health so ChatGPT can personalize health conversations. Connected data is not used to train models or target ads. GPT-5.6 Sol powers the experience for paid users, while GPT-5.5 Instant serves free users, both developed with input from hundreds of physicians worldwide.
Andrew Ng and Rohit Prasad have launched OpenWorker, an open-source AI agent for Mac that delivers finished work rather than just chatting. It can prepare customer briefs, draft reports, send Slack messages, and update calendar entries, checking in before taking consequential actions. OpenWorker supports multiple models including GPT, Claude, Gemini, DeepSeek, and Ollama for local inference. Users bring their own API key, and data stays on-device except through chosen providers. Windows support is coming soon.
Free subscriber bonus
Get the AI Income Database, free when you join
38+ real ways people are making side-hustle money with AI right now. Each one comes with the playbook and the exact tools to pull it off. No fluff and no course to buy. It's the bonus every new subscriber gets on day one.
- 38+ proven AI side-hustles, from first dollar to scale
- The exact tools for each one, picked from the 4,500+ we track
- Updated as new opportunities appear (subscribers hear first)
Joins the twice-weekly AI briefing read by 250,000+ people. Free forever, unsubscribe anytime.












































