MDLMPE: Distribution Aware Positional Encoding for Masked Diffusion Language Models
MDLMPE: distribution-aware positional encoding for masked diffusion language models to handle dynamic token-availability patterns during denoising.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
MDLMPE: distribution-aware positional encoding for masked diffusion language models to handle dynamic token-availability patterns during denoising.
GDPevo: benchmark for evaluating agent self-evolution on real business workflows with automated data pipeline to prevent contamination.
Faithfulness-safety tension in Large Reasoning Models: models must be interpretable for monitoring yet robust against unsafe reasoning paths.
Multi-agent clinical committees using Gemini show vulnerability to social shortcut cascades where peer consensus propagates errors across agents.
LLM capability for testing Terminal User Interfaces: benchmark across ratatui/Rust, bubbletea/Go, textual/Python showing 12% coverage in real applications.
Narrative review of AI-based sound effect generation across input modalities: text, visual, audio, and multimodal for digital applications.
MissClick: adversarial attack on GUI grounding models exploiting digit-serialized coordinate generation to induce large spatial click displacement.
AgenticECO: tool-using agent workflow for 3D-IC engineering change orders with minimal-disturbance routing layer on TaiWei open-source flow.
Taxonomy of multilingual multi-agent planning failures: request-to-action grounding degradation strongest in low-resource languages.
Self-augmentation method for MLLMs using model failure signals to generate targeted image augmentations without external supervision.
CARE-Bench evaluates 11 LLMs on medical triage safety, testing when patient-facing models recommend escalation to clinicians.
GPTKB 2.0 constructs disambiguated knowledge bases directly from LLM outputs using on-the-fly entity/relation resolution.
SAT-Edge-Agent deploys edge-based LLM agents on satellites for onboard intelligence under communication/power constraints.
Black-box diagnostic for LLM collectives measuring whether output diversity correlates with genuine epistemic revision.
Task system for detecting hallucinations in Arabic Islamic QA and selecting verified answers from candidates.
Amortized causal forecasting model for Cox-Ingersoll-Ross financial time series to estimate interventional outcomes.
Empirical study across 13 models showing letter casing modulates attention allocation in LLMs and VLMs.
Early-epoch telemetry from training runs predicts final test accuracy and training failure without cross-run reference.
Category-theoretic perspective on classical statistical learning models and algorithms for expository survey.
Competition-aware request dispatch for RTB ad exchanges using bid prediction and probabilistic forwarding to optimize DSP participation.
LiLa-WAM: lightweight latent-space world-action model for robotic manipulation with reduced computational overhead via single-stage training.
AntiSkillBench: end-to-end benchmark evaluating privacy leakage, impersonation risk, and defenses in persona-skill pipelines for agents.
Apple says its trade secrets investigation into OpenAI has widened. In a new court filing, Apple claims additional former staff may have retained or accessed confidential information.
TARL: memory state framework for long-term agents mapping statements to five executable actions (add, ignore, revise, reject, defer) instead of binary write/hold.
Learning and clustering on temporal graphs comparing GNN performance against classical algorithms for coarse-grained node aggregation.
GPU-accelerated community detection in dynamic graphs using NVIDIA RAPIDS with spectral clustering and Leiden optimization achieving 1000x speedup.
Pattern Completion Bias benchmark measuring how repeated UI patterns degrade multimodal LLM accuracy on screenshot-to-code fill-in-the-blank tasks.
LAEF: 7M-parameter lead-agnostic ECG foundation model processing variable lead subsets via spatiotemporal graphs for point-of-care diagnostics.
LiveEvalBench: agentic, adaptive evaluation framework treating web generation as interactive problem with diverse valid implementations.