Introducing ChatGPT Images 2.5
OpenAI releases ChatGPT Images 2.5 with two API variants (Sunburst, Flare) offering improved multi-turn instruction-following, faster generation, and better subject preservation in reference images.
Every story tagged with this topic, ordered by date.
OpenAI releases ChatGPT Images 2.5 with two API variants (Sunburst, Flare) offering improved multi-turn instruction-following, faster generation, and better subject preservation in reference images.
TANGO: vision-language-action model for humanoid robot navigation in cluttered indoor spaces with whole-body adaptation.
Study of image tokenizer design via multimodal continual pretraining on text, image, T2I, and I2T tasks.
NOAH: generative model for longitudinal multimodal patient data with irregular temporal dynamics and forecasting.
Probing study reveals vision encoders embed canonical color concepts linearly-decodable from grayscale images, demonstrating implicit semantic knowledge.
DeCAL integrates contact-aware latent co-imagination into vision-language-action models to handle dexterous manipulation with severe visual occlusions and contact dynamics.
GoDeep performs open-vocabulary 3D scene segmentation using vision-language models as translators without 3D training data or domain-specific encoders.
Gander: multimodal streaming agent supporting full-duplex video/speech/text interaction with real-time interrupts and proactive feedback.
AuK, open-source multimodal model for unified speech generation and editing via natural-language instructions and audio context.
Methodological comparison of MAE, Solar, and UniMMQA frameworks tracing evolution of multimodal QA architectures.
Kairos: fine-grained video-language dataset with time-resolved annotations across 10–30 min videos for temporal dynamics modeling.
OpenAI releases ChatGPT Images 2.5, an image generation feature enabling sketch-to-image and reference-guided personalization.
O2C-Nav enables zero-shot vision-language navigation in continuous environments with single MLLM call per step via spatial-aware waypoint generation.
DualRead separates answer capability from confidence calibration in medical vision-language models trained via GRPO.
ViS-CoT combines visual search with chain-of-thought reasoning for product attribute extraction from videos without fine-tuning.
CAROT aligns LLM cross-lingual representations via token-level optimal transport in language-agnostic space, improving multilingual transfer by isolating language-specific information.
OpenAI releases GPT-6 Astra with improved prompt understanding and 3D model generation capabilities for developers.
SignDino self-supervised encoder adapts DINOv3 from spatial to temporal domain for sign language video representation learning.
RoboSPA benchmark evaluates Vision-Language-Action models on spatial reasoning and procedural complexity in robotic manipulation tasks.
β-VAE ECG representations from ECGx.AI foundation model discriminate myocardial scarring patterns for cardiac diagnosis screening.
SMILE framework enables self-explainable multimodal medical diagnosis via information bottleneck optimization.
Layer-wise causal intervention analysis of vision-language models reveals visual information integration during candidate processing.
SciDocBench: workflow-centered benchmark with 124 expert questions across scientific domains testing joint reasoning over text, equations, figures, tables, code, and datasets.
Greg Brockman discusses OpenAI's Astra multimodal model, organizational history, and alignment challenges in wide-ranging Stratechery interview.
Seeing Before Synthesizing uses VLM-generated visually grounded captions instead of LLM-only synthesis for weakly-supervised dense video captioning with adaptive linguistic guidance.
AdaRoboVLG decouples foundation models from grasp policies via kinematic mapping and force-closure stability, enabling generalizable vision-language grasping across robotic hands.
CORE: distill reranker compositional reasoning into MLLM embeddings via Rank-KL objective for scene understanding.
Editable Visual Design combines VLM creative reasoning with Coding Agent precision to generate layered, editable visual designs beyond flattened diffusion outputs.
InSituMeasure benchmark evaluates MLLMs on continuous-valued measurement tasks in realistic industrial settings with gauge reading and instrument-specific context.
Text-AB unified diffusion-transformer framework for voice dubbing and full-duplex dialogue via alignment-free latent synthesis with DAC-VAE compression.
Audit of zero-shot vision-language models (CLIP, BioCLIP, BioCLIP2, Jina CLIP v2) on Bangladeshi freshwater fish recognition revealing multilingual and nomenclature limitations.
DiscoSign adds discourse-level reasoning to LLM-based text-to-sign-language translation for spatial coreference and QACs.
ShallowStream optimizes streaming video MLLMs by pruning computation at shallow layers, reducing overhead for embodied AI and autonomous driving.
Training-free RVSD framework reduces visual hallucinations in vision-language models via retrieval and sparse decoding.
Pseudo-triplet construction pipeline enables direction-following text-to-speech generation without paired training data of script modifications and delivery directions.
Multimodal speech-language framework predicts loneliness in older adults from telephone interviews using psycholinguistic features and topic modeling.
VMetaphor-Bench: first benchmark for visual metaphor generation in text-to-image models; 1,500 curated samples across three levels and ten categories.
ViSAR uses training-free adaptive retrieval for Document VQA, dynamically selecting page count per query to reduce LVLM latency.
Retrieval-augmented generation and BioNER improve lay summarization of radiology reports to reduce patient reliance on general LLMs.
MultiGhostBench: 928-book multilingual benchmark for detecting LLM-generated text across six languages and distribution shifts.
Multimodal framework for subcellularly resolved cell embeddings using RNA, protein sequences, and structural information.
SonicCaps: 15M-caption audio dataset with fine-grained descriptions from Qwen3-Omni for improved audio-language retrieval.
H3-World framework leverages MiniMax-H3 video generator's emergent language interface for grounded world control without dedicated action modules via structured language prompting.
Google DeepMind extends Gemini with agentic video understanding capabilities for autonomous analysis and reasoning over video content.
TempCloze: 1,521-video benchmark for temporal reasoning in Video-LLMs, minimizing language shortcuts via cloze-completion on egocentric clips.
Google launches Pics, an image creation/editing tool built on Nano Banana model, integrated into Google Workspace.
Vision Transformer method for cancer grading combines nuclei classification with histopathology images via semantic-guided multimodal preprocessing.
DaEdiTikZ dataset leverages scientific revision trajectories to train VLMs for iterative scientific figure editing without synthetic supervision.
Contribution-aware bandwidth allocation optimizes split learning for multimodal edge training by allocating uplink budget based on modality importance.
DroneCATS benchmark evaluates MLLMs as generalist vision-language-action agents for drone control with full action-space prompting.