API Benchmark Scores Do Not Reliably Transfer to Chatbot Interfaces
API benchmark scores for ChatGPT, Claude, Gemini systematically diverge from chatbot interface behavior by 3.4pp average; questions reliability of model eval standards.
Every story matching this topic across titles and summaries, newest first.
API benchmark scores for ChatGPT, Claude, Gemini systematically diverge from chatbot interface behavior by 3.4pp average; questions reliability of model eval standards.
The sheriff’s office said the hikers “were advised by Gemini to bring far less food and water than their group required."
Study of substrate blindness in AI agents: Claude Opus 5, GPT-5.6-Sol, Gemini 3.7 Flash code generation ignoring memory/compute constraints.
Gemini Spark can edit and curate photo albums, create shared collections, turn photos into calendar events, and handle other Google Photos tasks for AI Pro and Ultra subscribers.
Service interruptions hit ChatGPT, Claude, Grok, and Gemini practically simultaneously.
Google is rolling out AI-powered voice assistant modes for Gmail, Docs, and Keep that allow you to manage the apps by talking to them. The real time conversational capabilities are called Gmail Live, Docs Live, and Keep Live, and like the Gemini Live experience for Google's chatbot, aim to make it easier to note down or find specific information when you're on the move or unable to do so manually. Gmail Live is primarily designed to surface details from your inbox without you having to dig through lengthy email threads using keywords or subject lines. For example, Google says that you can ask...
Google launched Gemini 3.8 Flash, arriving just a few weeks after its predecessor. The company claims the new model "works harder" than Gemini 3.7 Flash by performing more reasoning steps on complex tasks and "calling tools iteratively." It has the same introductory pricing as 3.7 Flash, $0.75 per million input tokens and $3.75 per million output tokens, but could still end up costing users more. Google warns that "the model might use more tokens to maximize performance, especially at higher effort levels." Developers can keep using Gemini 3.7 Flash if they want to minimize token usage. Gemin...
Google's Pro model updates are seemingly paused, but there's yet another Gemini Flash today.
llm-gemini 0.34 adds Gemini 3.8 Flash support with configurable thinking levels and async bug fixes.
Google DeepMind releases Gemini 3.8 Flash and specialized Gemini 3.8 Flash Cyber variant for security applications.
MrBeast will feature Gemini, Google Health, and the Fitbit Air in upcoming videos as part of a multi-year partnership with Google. The deal will kick off with a video featuring Jimmy "MrBeast" Donaldson turning to Gemini for wilderness survival advice: First up on September 5 is a new MrBeast video following Jimmy and his crew as they attempt to survive some of the most extreme and treacherous climates on Earth. Teams will race to navigate three of the planet's most unforgiving landscapes - the jungle, the desert, and the Arctic - using Gemini to help them identify dangers, survive brutal wea...
Google DeepMind extends Gemini with agentic video understanding capabilities for autonomous analysis and reasoning over video content.
Google wants Workspace users to edit and generate their business imagery with Pics. | Image: Google Google has a new suite of creative design tools for Workspace users called Google Pics, which aims to make editing and generating "professional-grade" AI images less cumbersome for businesses. Built around Gemini and the Nano Banana generative AI model, Google Pics is designed to give more granular control over prompt-based image making and manipulation, allowing users to tap on specific objects or text, and describe what they want to change - features that may avoid the ugly results that chatb...
Versions of OpenAI's ChatGPT and SpaceXAI's Grok will join Google's Gemini on the Pentagon's central portal for AI tools.
Google's AI note-taking app, Gemini Notebook, can now pull information from the books you've purchased. The new "Expert Intelligence" feature allows you to bring titles from Google Play Books directly into Gemini Notebook, which means you can ask questions about the material, as well as generate plans, infographics, AI podcasts, and more based on their contents. During a briefing with The Verge, Google Labs editorial director Steven Johnson showed how you can use the tool to generate a recipe book using the information in Michael Pollan's Food Rules. In another example, Gemini Notebook applie...
Google DeepMind releases Gemini 1.1 Flash with enhanced control features for developers building applications.
Consumer AI apps need to stop making users learn their product architecture.
The AI that powers Gboard's Rambler is coming to more Google products, including Chrome.
Google DeepMind launches Gemini 3.5 Transcribe, an upgraded speech-to-text model with improved contextual understanding.
Google has updated Gemini Audio with some new Gemini 3.5 models, introducing new transcription capabilities that automatically detect specialized jargon and more than 85 languages. Gemini 3.5 Live, 3.5 Live Experimental, and 3.5 Transcribe are designed to provide better precision for Google's voice-controlled AI features, without struggling with background noise or when your speech is interrupted. Gemini 3.5 Transcribe is a completely new addition to the Gemini family, and its introduction comes as we're still waiting for Google to release the Gemini 3.5 Pro model that it promised to roll out...
Cross-agent specification portability: Claude, Gemini, Copilot on Oracle-to-PostgreSQL migration; 380/1006 successful executions.
The launch of the new study features marks Google's latest effort to make Gemini the AI assistant that students turn to when learning and studying, as it continues to compete with companies like OpenAI.
As we're gearing up for back-to-school season, Google is rolling out a new dedicated student hub in Gemini. It's a one-stop repository for collecting research in a study notebook, creating flashcards, taking practice quizzes, and more. Google is also enhancing its study notebooks with support for graphs and images. It can even add test dates and deadlines to your Google Calendar based on your syllabus. Google is also adding Deep Research to Gemini Live. You can ask Gemini to generate complicated research reports and talk through the results. If it's taking a while, you can close the chat and ...
One of the best things my smart home does is help me care for my pets, and security cameras are particularly useful for keeping track of my many critters. But the barrage of notifications they send often means I miss important ones. So, when Google announced its new Pet Memory feature for Gemini for Home, I thought this could be the perfect solution. Pet Memory promises to teach your Google Home smart home who your pets are, allowing connected Nest cameras to go beyond simply telling you they saw an animal to which animal they saw. It can then use that knowledge to adapt your smart home to yo...
Google Gemini and Pixel partner with five global football clubs to elevate the fan matchday experience through AI and Smartphone Technology.
Google will now allow you to remove visible watermarks from the images, videos, and music made with AI tools. With the update, you can toggle off a new "Media watermark" setting in Gemini and Google's AI video generator, Flow. When toggled off, Google will remove the "sparkle" watermark that appears in the bottom-right corner of content generated with the company's Nano Banana and Omni models. Though visible watermarks are now optional, AI-generated content will have invisible SynthID watermarks and C2PA metadata embedded in the background, according to Josh Woodward, the vice president of Go...
Release: llm-gemini 0.33 It's been a while since the last llm-gemini release. This version of the plugin adds support for today's Gemini 3.7 Flash release, plus gemini-3.6-flash , gemini-3.5-flash-lite and two embedding models gemini-embedding-2 and gemini-embedding-001 . The plugin is also upgraded for compatibility with LLM 0.32, which means you can now see reasoning traces and you can also enable server-side tools using this pattern: llm -m gemini-3.7-flash -T CodeExecution \ 'use python to calculate (factorial of 13) * 3' I had Gemini 3.7 Flash draw me some pelicans riding bicycles at hig...
Gemini 3.6 Flash debuted just 3 weeks ago, but Google says 3.7 has "substantial improvements."
From the Pixel 11 series and a brand new competitor to Apple’s AirTag, here are all the announcements from the Made by Google 2026 event.
But will Gemini's surge survive slowing model releases?
Gemini is keeping pace with OpenAI’s ChatGPT, which hit 1 billion monthly active users back in June.
Assistant will disappear, leaving only Gemini for voice control in the coming weeks.
Farewell, Google Assistant art, we hardly knew ye. | Illustration by Alex Castro / The Verge Google Assistant's days have been numbered ever since Gemini arrived on the scene, and its time is now up. Google has announced that it will be removing access to Assistant on Android phones and tablets, along with paired devices like smartwatches or headphones, from September 4th. The announcement came in an email apparently sent to some users, as reported by 9to5Google and shared in full on Reddit. The Verge has asked Google for confirmation that the email is accurate, and will update this story onc...
Multi-agent clinical committees using Gemini show vulnerability to social shortcut cascades where peer consensus propagates errors across agents.
Study compares GPT-5.4, Gemini 3.1 Pro, and Claude Opus 4.6 peer reviews on 300 ICLR submissions against human reviewer alignment.
Microsecond-cost anomaly detection for LLM agent failures using one-class echo-state networks trained only on healthy runs; tested across Qwen, Llama, and Gemini agents.
Now a text prompt is all it takes to generate reality-warping images using Google Earth's satellite, aerial, and 3D imagery, like these images generated by Digital Digging's Henk van Ess that show "refugees near the Mexican border" and a bomb crater near a hospital in Gaza. Google responded to Digital Digging's AI-altered images, saying, "We take misinformation seriously - every image created with Nano Banana in Google Earth includes the SynthID digital watermark, so if someone is unsure about an image, they can ask the Gemini app or use Lens in Search to see if the image was AI-generated." I...
Gemini Robotics 2 includes three models, but only one is publicly available right now.
Apptronik’s Apollo 2 robot takes a baseball glove off of a shelf. | Image: Google Google DeepMind says the latest version of its Gemini Robotics AI model can "control entire humanoid robots." While the previous model focused on controlling a humanoid robot's upper body, Gemini Robotics 2 now supports "whole-body motions" ranging from its feet to fingertips, according to an announcement on Thursday. The new model will allow humanoid robots to perform a wider range of actions, as it allows them to walk, crouch, stretch, and manipulate objects. Videos shared by Google show how Apptronik's Apollo...
Google DeepMind releases Gemini Robotics ER 2 with video understanding, task orchestration, and multi-robot coordination capabilities.
Claude Opus-4.7, GPT-5.4, Gemini-3.1-Pro confabulate medical diagnoses without images; diagnosis systematically shifts by patient demographic, raising safety concerns.
Benchmark of six VLMs (Gemini, GPT-4V, Qwen, Gemma, Llama, Ministral) on zero-shot anomaly detection for game geometry clipping in agent-driven QA.
Google ships Gemini API Managed Agents with Gemini 3.6 Flash and tool-use hooks for production agent deployment.
Google DeepMind releases Gemini Robotics 2, extending multimodal foundation model capabilities to whole-body robot control and coordination.
Controlled study compares LLM (ChatGPT-4o, Gemini) vs. human literature review performance across physics, astrophysics, cosmology.
Willison surveys evolving AI tool recommendations, noting shift from chat interfaces to agentic systems; Gemini absent from current guidance due to lack of autonomous work capability.
Study finds Grok assigns 2-5x higher credibility to ethnonationalist pseudo-science than Claude, GPT, Gemini across four LLM families.
Meta says its AI chatbot is going beyond just answering questions and generating images. | Image: Meta Meta is upgrading its AI chatbot with new productivity features in a bid to compete with rivals like Gemini, ChatGPT, and Claude. The update will allow Meta AI to tap into your calendar to help you plan events and generate daily briefings, as well as perform in-depth research that you can steer as it progresses. In a blog post, Meta says this update marks its "next step toward personal superintelligence," something CEO Mark Zuckerberg has touted as the future of AI. Meta is powering the upda...
Black Forest Labs releases FLUX 3 multimodal model with reported improvements over Gemini 2.0, Grok Imagine, and includes video-action robotics variant.
Gemini had over 750 million monthly users in February.
Web analytics (Aug 2023–Oct 2025) show ChatGPT, Perplexity, Gemini drive measurable referral traffic to academic library resources, particularly theses databases.
There are new 3.6 and 3.5 models today, but Google is already training Gemini 4.
Google DeepMind releases Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber models for inference.
Google is launching Gemini 3.6 Flash alongside a new security model dedicated to quickly finding and patching security vulnerabilities. In a blog post on Tuesday, Google describes Gemini 3.5 Flash Cyber as a "cost-efficient and highly capable alternative" to larger, more expensive AI systems, such as the one offered by Anthropic's Mythos. The cybersecurity model is built upon Gemini 3.5 Flash and will be available first to governments and trusted partners via CodeMender, Google's security-focused coding agent. As noted by Google, CodeMender can call upon 3.5 Flash Cyber "multiple times at hig...
Alphabet, Google's parent company, is reportedly working on a new chip designed to make its Gemini models run much more efficiently.
Domain-generalized pixel-level tampering detection robust across VLM-generated manipulations from ChatGPT, Gemini, Qwen-Image.
Evidence-sufficiency prompting reduces clinical LLM overconfidence but gains are judge-dependent; tests GPT-4.5, Claude Opus, Gemini, Grok on real data.
Google DeepMind introduces Gemini 3.5 Flash Cyber, a specialized model for vulnerability detection and patching.
Google is adding personalized AI avatars to Vids that let users create videos starring a digital version of themselves, alongside Gemini Omni-powered tools for generating and editing videos from prompts and reference images.
Google is giving its AI note-taking app a new name. The company announced on Thursday that NotebookLM is becoming Gemini Notebook, but will remain a standalone app even as it integrates more deeply across Gemini and Google Search. Google first revealed Gemini Notebook - then called Project Tailwind - in May 2023 before widely releasing the app just months later. Over the past few years, Google has been adding new features to the app to help organize and make sense of your notes, such as the ability to summarize them as AI podcasts, narrated slideshows, and TikTok-style clips. It recently star...
Google Vids adds Gemini Omni support and personal avatar features for video generation and editing.
Google must give rival AI assistants and search engines greater access to key parts of Android and Google Search after the European Union ordered the company to comply with the bloc's digital antitrust rules. The two decisions, handed down Thursday, could weaken Google's control over two of the tech industry's most important platforms and have far-reaching consequences for the company, shape the future of its AI tool Gemini, and open up new opportunities for rivals to gain ground. Google has until January 2027 to begin sharing search data and July 2027 to implement changes to Android. The rul...
Google DeepMind and AIM launch ATL Saathi, a Gemini-powered educational tool for Indian robotics labs.
Waze is getting an AI makeover. Google is integrating its flagship AI assistant, Gemini, into the driving app with the goal of letting users personalize their trips a little more. Of the four new updates, only two are being described as involving Gemini. Waze says its updating its conversation reporting feature, first introduced in 2024, to allow drivers to use conversational voice commands to report traffic incidents and suggest map updates, like a road closure or outdated house number. In addition, Waze introduced Destination Search, enabling drivers to use (again) use conversation voice co...
Android Bench is evolving, and developers can help guide that process.
Google expands Gemini API Managed Agents with background task execution and remote MCP support for production deployments.
I call BS: the founding fathers definitely would have been Microsoft Teams users. | Image: Google "Group project, but make it 1776." That's how a new commercial for Google Workspace opens. And things only get cringier from there. The clip imagines what it would be like if the founding fathers turned to Google's collaboration tools and Gemini to help them draft the Declaration of Independence. Ben Franklin texts Thomas Jefferson to check on the status of a draft, who takes a photo and uses AI to transcribe it into a Google Doc. Franklin and Adams hop in to make edits in suggestion mode, Gemini...
Comparative evaluation of frontier LLMs (GPT, Claude Opus, Gemini, GLM) for automated Linux/bash exam grading using cognitive taxonomy.
Multi-agent workflow (Gemini 2.5 Flash, RigoChat-7B-v2) for Spanish Easy-to-Read translation via LangGraph with event-condition-action routing.
MedQADE benchmark reveals LLM evaluators (Gemini 3 Flash) match clinician agreement on German medical QA but lack clinical caution in safety assessment.
Let’s start with a game. Open up your chatbot of choice—Claude, ChatGPT, Gemini—and type “Give me a random number between 1 and 10.” You’re going to get 7. Almost always. Now type “Another” and you’ll get 3 or 4. Type “Another” again and you’ll get 8 or 9. That won’t work every time—but if it…
Google's 24/7 agentic assistant, Gemini Spark, comes to Mac alongside other improvements, like real-time tracking and support for more apps.
Zero-shot benchmark evaluates Claude, GPT-5.4, and Gemini on fine-grained 13-class emotion classification.
The Google Home Speaker is Google’s first smart speaker in years. And it’s pretty! | Photo: Jennifer Patison Tuohy / The Verge Smart speakers have spent the past few years searching for a compelling second act. Beyond music, timers, and controlling your lights, they've struggled to justify taking up space on the kitchen counter. AI promised to change that. Amazon debuted its new hardware powered by a revamped Alexa last fall, and now it's finally Google's turn. The Google Home Speaker is the company's first new smart speaker in six years and its first "built for Gemini." After years of neglec...
Google releases Gemini 3.1 Flash Lite, optimized for fast, low-cost image generation; author tests visual search capability.
Anthropic’s Claude Sonnet 5 brings stronger agentic capabilities, lower pricing, and improved safety, positioning the model as a cheaper alternative to Opus, GPT-5.5, and Gemini Pro.
Google DeepMind releases Gemini Omni Flash and Nano Banana 2 Lite for developer access.
Google is expanding Gemini’s personalized AI image generation to eligible free users in the U.S., allowing the chatbot to create images based on your interests and data from connected Google apps.
Speech conveys information through both words and vocal delivery. We evaluate four leading production realtime voice systems-OpenAI's GPT Realtime 2, Google's Gemini 3.1 Flash Live, and Alibaba's Qwen3.5 Omni Plus and Omni Flash-on tasks where the words and the delivery patterns both convey meaningful information. Across three consequential scenarios, all four systems act on the words rather than the voice. They end calls with crying callers who insist nothing is wrong, approve wire transfers authorized in frightened voices, and enroll callers whose agreement is clearly sarcastic. Surprisingl...
Here's what you need to do to get those pesky "write with Gemini" pop-ups to go away.
Google is betting generative AI can breathe new life into the smart speaker. The company's new $99.99 Google Home Speaker replaces the rigid commands of the Google Assistant era with more conversational Gemini interactions.
Google's new smart speaker is more about Gemini than audio quality.
Google has released Android 17 and Wear OS 7, introducing new multitasking features, parental controls, security tools, and smartwatch upgrades. The launch is also accompanied by a Pixel Drop that brings Google’s latest AI models to its devices.
The chatbot still remains the most popular AI assistant worldwide with over 1.1 billion monthly users, followed by Gemini with 662 million and Claude with 245 million.
When I returned to my computer five minutes after giving Gemini a lengthy prompt, I had two things: a functional app in a preview window, and a message about a bug. "~ Channel is unrecoverably broken and will be disposed!" Sounded bad! But right below it was a button to fix the bug. Pretty weird that I just instructed a computer to build a whole app for me with a single prompt, but it needed me to click a button to fix a bug. I did anyway, and in 233 seconds Gemini reported back that it had succeeded, using words like "blockages" and "race conditions." I didn't understand a bit of it. It was ...
The fraudsters allegedly targeted hundreds of thousands of people with Gemini-coded scams sites.
DiffusionGemma Last May Google briefly released an experimental Gemini Diffusion model. I tried the preview at the time and recorded it running at 857 tokens/second. It was an exciting model, but Google made no further announcements about it. That research has returned in the best possible way: as a new open weight (Apache 2 licensed) Gemma model, google/diffusiongemma-26B-A4B-it . NVIDIA are currently hosting the model for free on their NIM cloud API. I used that API to generate this pelican , which took 4.4s (according to time uv run generate.py ) to return 2,409 tokens - so at least 500 to...
This paper presents our solution to the 2026 SoccerNet VQA Challenge. We first develop a cost-effective data synthesis pipeline driven by a Vision-Language Model (VLM), which systematically restructures raw domain data into diverse VQA samples, including concise answers and long-form responses. Second, we propose MSUE, a multi-expert question answering architecture that employs a Large Language Model (LLM) to dynamically dispatch questions to text, image, and video experts. These experts are instantiated as a strong text baseline Gemini3-Flash, a fine-tuned Qwen3-VL, and an external knowledge...
Voice translations preserve speaker's tone, pacing, pitch—with SynthID watermarks for security.
This study investigates cross-lingual distributional skew (the Shibboleth Effect) in frontier large language models (LLMs) subjected to sustained adversarial conditions. We develop a multi-agent geopolitical wargame, the Cerulean Sea Crisis, a synthetic maritime territorial dispute designed to mirror the structural dynamics of Eastern Mediterranean conflicts. Six frontier models (GPT-4o, Llama-4, Mistral-Large, Gemini-3.1-Pro, Qwen3.6-Plus, and DeepSeek-R1) participate in a between-groups experiment (N = 10 games per arm, K = 5 rounds per game) in which the sole manipulation is the language o...
Large language models are increasingly used to adapt math word problems for personalized learning at scale, but it remains an open question whether those adaptations are consistent across models, preserve cultural diversity at scale, and reveal which cultural entities models treat as most salient. We analyze how Claude Opus 4, GPT-4.1, and Gemini 2.5 Pro adapt 60 English math word problems into Bengali, Hindi, Punjabi (India), Urdu, Sindhi (Pakistan), Italian, and Sicilian (Italy), a language set spanning the full resource spectrum, from high-resource Italian and Hindi to under-studied Sindhi...
Gemini 3.5 Live Translate brings near real-time, natural speech translation to Google AI Studio, Google Translate and Google Meet.
Given how badly burned anyone who took Apple's 2024 WWDC Apple Intelligence announcements at face value was, I'm holding to a strict "I'll believe it when I see it" policy for everything they announced today . The new Siri AI features do at least look feasible with today's technology, especially since Apple are licensing a custom Gemini-derived model that they can run on their own Private Cloud Compute. It sounds like they'll be taking advantage of vision-LLMs to extract information from the user's screen, which neatly sidesteps the need for every existing application to ship custom code in o...
NotebookLM is getting a big upgrade, but it's only for AI Ultra and enterprise accounts right now.
Google is rolling out "across the board" updates to NotebookLM. The AI-powered note-taking app now uses Google's upgraded Gemini 3.5 model, which will allow it to respond with "more accurate and reliable information," according to a blog post on Monday. Launched in 2023, NotebookLM allows you to interact with your notes and sources using AI, as well as ask questions about the materials. With this update, Google says you can start a research project by just asking NotebookLM questions about a topic, instead of importing notes or YouTube videos. NotebookLM will use Google Search to help you fin...
Results from a randomized controlled trial show the potential of Gemini’s Guided Learning feature to boost engagement and accelerate learning.
This week we've got tandem hands-ons with Google's new Gemini AI agent - Spark - from my colleagues David Pierce and Jay Peters. Their takeaways are similar: It's so effective that it's scary. Spark knew that David's dog is named Frida and knew the first name of Jay's wife, even though neither of them explicitly provided this information to Google. But what's scary to me is how all of this stuff seems geared toward a future of "productivity" that completely misses what needs to be fixed in our world. "Productivity" is often pitched as a panacea for what befalls us in our personal lives, even ...