3D-Aware VLMs with Implicit and Explicit Geometries
VLM-IE3D framework enhances vision-language models with implicit and explicit 3D geometry tokens from RGB video for improved spatial reasoning.
Fresh arrivals from arXiv cs.AI, cs.CL, and cs.LG. The raw research feed.
VLM-IE3D framework enhances vision-language models with implicit and explicit 3D geometry tokens from RGB video for improved spatial reasoning.
Expanding Flow Maps (EFMs) enable flow-based generative models to handle variable-dimensionality distributions via expanding interpolants with conditional noise.
GraphVid enables precise multi-object video generation control via graph-structured representations instead of trajectory or text constraints.
Theoretical analysis proves Barzilai-Borwein optimization method fails superlinear convergence on open set of quadratics for dimension n≥4.
Synthetic data generation framework using deep learning to automate surface defect detection in rotogravure printing quality control.
Philosophical critique: surprisal theory's linguistic difficulty predictions are tautological without constraints on language model specification.
TimePNS framework for time-series model explanation using counterfactual necessity to identify essential (not spurious) decision factors.
MedGame transforms static clinical cases into interactive decision-driven learning games using LLMs and dual narrative/director engines.
EnsembleEGNN molecular foundation model encodes conformational ensembles of cyclic peptides using equivariant GNNs with set attention pooling.
Consensus anomaly detection applied to Ghana malaria surveillance data identifies spatiotemporal hotspots in Ashanti, Northern regions 2014-2023.
Study reveals LLM moral reasoning involves structured resistance-compliance dynamics paralleling human social psychology, beyond simple sycophancy reduction.
OpenForgeRL enables end-to-end training of harness-native agents with open infrastructure, addressing limitation of complex inference harnesses like Claude Code.
VCSD proposes visual contrastive self-distillation removing need for privileged information in on-policy distillation via pure input conditioning.
MIRROR framework exploits complementary reasoning paths across text, diagram, and combined modalities to improve vision-language model reasoning on geometry problems.
X³-OPD cross-modal distillation framework transfers reasoning from text LLM teacher to audio-language student via on-policy alignment and acoustic perception.
Neural networks solve coupled Dyson-Schwinger equations for Yang-Mills gauge theory with percent-level agreement to fixed-point solutions.
Theory paper argues human participation persists in automated systems for technical, complementarity, and normative reasons beyond current AI capability limits.
Zero-Flow Two-Sample Test uses learned directional misalignment patterns for distribution testing, separating witness learning from hypothesis evaluation.
DONDO releases 26 open w2v-BERT speech recognition models for African languages spanning six countries, trained on religious text corpora.
Windowed-MTP optimizes speculative decoding at million-token context by eliminating full-KV attention overhead in multi-token prediction draft heads.
Petri-net-guided LLM test generation for concurrent Rust APIs addresses shallow test synthesis by integrating formal models with executable test concretization.
ElasticTTT framework prevents prior collapse in test-time tuning of diffusion models for video editing by preserving distribution-mapping during optimization.
GS-Agent generates physically plausible 4D worlds from natural language by combining foundation models with agentic simulation and physics constraints.
Study using gpt-5.6-sol shows LLMs produce safer advice when dangerous objectives are mediated through agent transformation versus direct exposure.
Improved lower bounds for Shannon capacity of odd cycles via independent set construction in graph powers—pure graph theory unrelated to AI.
Agentic context management frames token cost and memory bloat as lifecycle and architecture problems, not storage-retrieval, for production agent reliability.
LLMs systematically overuse epanorthosis (classical self-correction rhetoric) due to promotional training distributions and RLHF preference for emphatic phrasing.
Speech-based multimodal LLMs detect cognitive impairment across diverse speakers and devices by leveraging linguistic and acoustic biomarkers with improved generalization.
No-code agent platforms create reliability gaps—silent degradation from changing models, tools, permissions, and dependencies—requiring continuous assurance frameworks.
Analysis of code model representations shows Qwen2.5-Coder and DeepSeek-Coder align on grammatical concepts across Python/Rust, with task-driven specialization.
MAPS: hierarchical MARL system using centralized proto-plan embeddings for decentralized AV coordination at unsignalized intersections.
Open-source evaluation framework for open-weight LLM agents on longitudinal data tasks, addressing privacy constraints in research deployments.
Label complexity bounds for auditing high-recall candidate generation pipelines with finite-sample validity guarantees.
Randomized KV-cache eviction with error certification via Hájek correction, proving deterministic eviction hides information loss.
Thinkink: 2D spatial interface integrating handwritten/sketch prompts with LLM responses via semantic tree interpretation.
AREX: recursively self-improving research agent exploiting discovery-verification asymmetry to refine multi-constraint answers.
Token-level detection method for LLM-generated content in human-AI coauthored text using score smoothing.
TTEL: inference-time algorithm using token-level error localization and environment feedback for efficient test-time scaling.
RUMBA: Russian benchmark for long-term LLM conversational memory with fine-grained taxonomy across temporal reasoning dimensions.
KroQuant: Kronecker-structured block transforms for W4A4 post-training quantization of diffusion transformers with efficient inference.
TriviaRoomQA benchmark evaluates multilingual LLM performance on 3,300 culturally-grounded trivia questions across 6 European languages and long-tail knowledge.
FGDSE framework applies causal-ensemble methods to predict EV charging infrastructure faults under climate stress for preventive maintenance.
Concept-based agent-guided learning improves interpretability and generalization of deep learning models for surgical margin assessment via REIMS spectroscopy.
Adaptive Identity Anchoring improves video face swapping by optimizing keyframe placement for synthetic paired supervision in identity transfer.
Linear probes on hidden states detect early non-convergence in chain-of-thought reasoning; DeepSeek-R1-Distill-Qwen-7B shows 90.3% converged vs 6.6% non-converged AIME accuracy.
Context-weighted Discrete Flow Matching modifies CTMC to weight training targets by local context density, improving generative modeling on discrete structures.
Semantic-aware task clustering for Cooperative Multi-Task Semantic Communication (CMT-SemCom) ensures constructive multi-tasking by aligning tasks post-initialization.
Multi-axis evaluation framework for structured audio captions on AudioCards dataset validates five orthogonal dimensions beyond flat text metrics.
Constraint-aware flow maps apply symbolic filtering, weighting, and repair to conditional diffusion models for dynamically feasible graph trajectory generation.
PATS reframes skills as dynamic training scaffolds for LLM agent reinforcement learning, converting rollout groups to reduce failure repetition in long-horizon tasks.