3D-Aware VLMs with Implicit and Explicit Geometries
VLM-IE3D framework enhances vision-language models with implicit and explicit 3D geometry tokens from RGB video for improved spatial reasoning.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
VLM-IE3D framework enhances vision-language models with implicit and explicit 3D geometry tokens from RGB video for improved spatial reasoning.
Expanding Flow Maps (EFMs) enable flow-based generative models to handle variable-dimensionality distributions via expanding interpolants with conditional noise.
GraphVid enables precise multi-object video generation control via graph-structured representations instead of trajectory or text constraints.
Theoretical analysis proves Barzilai-Borwein optimization method fails superlinear convergence on open set of quadratics for dimension n≥4.
Synthetic data generation framework using deep learning to automate surface defect detection in rotogravure printing quality control.
Philosophical critique: surprisal theory's linguistic difficulty predictions are tautological without constraints on language model specification.
TimePNS framework for time-series model explanation using counterfactual necessity to identify essential (not spurious) decision factors.
MedGame transforms static clinical cases into interactive decision-driven learning games using LLMs and dual narrative/director engines.
EnsembleEGNN molecular foundation model encodes conformational ensembles of cyclic peptides using equivariant GNNs with set attention pooling.
Consensus anomaly detection applied to Ghana malaria surveillance data identifies spatiotemporal hotspots in Ashanti, Northern regions 2014-2023.
Study reveals LLM moral reasoning involves structured resistance-compliance dynamics paralleling human social psychology, beyond simple sycophancy reduction.
OpenForgeRL enables end-to-end training of harness-native agents with open infrastructure, addressing limitation of complex inference harnesses like Claude Code.
VCSD proposes visual contrastive self-distillation removing need for privileged information in on-policy distillation via pure input conditioning.
MIRROR framework exploits complementary reasoning paths across text, diagram, and combined modalities to improve vision-language model reasoning on geometry problems.
X³-OPD cross-modal distillation framework transfers reasoning from text LLM teacher to audio-language student via on-policy alignment and acoustic perception.
Neural networks solve coupled Dyson-Schwinger equations for Yang-Mills gauge theory with percent-level agreement to fixed-point solutions.
Theory paper argues human participation persists in automated systems for technical, complementarity, and normative reasons beyond current AI capability limits.
Zero-Flow Two-Sample Test uses learned directional misalignment patterns for distribution testing, separating witness learning from hypothesis evaluation.
DONDO releases 26 open w2v-BERT speech recognition models for African languages spanning six countries, trained on religious text corpora.
Windowed-MTP optimizes speculative decoding at million-token context by eliminating full-KV attention overhead in multi-token prediction draft heads.
Petri-net-guided LLM test generation for concurrent Rust APIs addresses shallow test synthesis by integrating formal models with executable test concretization.
ElasticTTT framework prevents prior collapse in test-time tuning of diffusion models for video editing by preserving distribution-mapping during optimization.
Runway no longer wants to be just another AI model company. It wants to become the infrastructure layer for generative media. On Thursday, the startup launched Runway Media Router through Runway Dev, its developer platform, released earlier this month, that provides API access to a growing roster of third-party image, video and audio models alongside […]
GS-Agent generates physically plausible 4D worlds from natural language by combining foundation models with agentic simulation and physics constraints.
Study using gpt-5.6-sol shows LLMs produce safer advice when dangerous objectives are mediated through agent transformation versus direct exposure.
Improved lower bounds for Shannon capacity of odd cycles via independent set construction in graph powers—pure graph theory unrelated to AI.
Users can also integrate their personal data from services like Apple Health, Function, and MyFitnessPal.
OpenAI is rolling out ChatGPT Health to everyone in the US on Thursday, allowing more people to connect their medical records and health-tracking information to the chatbot. During a briefing, Ashley Alexander, OpenAI's vice president of health product, says the company's models "are now capable of reasoning at levels that are better than clinician level." When asked for more information about how the performance of OpenAI's models stacks up against human clinicians, OpenAI health lead Karan Singhal says that he would "temper" the claim that they are reasoning at better levels, but that "ther...
Agentic context management frames token cost and memory bloat as lifecycle and architecture problems, not storage-retrieval, for production agent reliability.
LLMs systematically overuse epanorthosis (classical self-correction rhetoric) due to promotional training distributions and RLHF preference for emphatic phrasing.