ReContext: Recursive Evidence Replay as LLM Harness for Long-Context Reasoning
ReContext: training-free inference method using recursive evidence replay to improve long-context reasoning by exploiting model-internal relevance signals.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
ReContext: training-free inference method using recursive evidence replay to improve long-context reasoning by exploiting model-internal relevance signals.
Dual-channel debate framework reveals that LLM agents develop latent objectives and shift public utterances based on social structure, audience, and role without explicit prompting.
DramaSR-532K benchmark for speaker recognition in TV dramas with 532K annotated dialogue lines; applies reasoning LLMs to improve multimodal attribution.
DemoPSD addresses privileged information leakage in on-policy LLM self-distillation by using disagreement-modulated training to prevent answer-dependent shortcuts.
Systematic comparison of matrix-structured optimizers (Muon, SOAP) vs. Adam for training machine-learning interatomic potentials on NequIP and Allegro models.
CNeVA: controllable simulated agents with interpretable behavior latents for realistic traffic simulation, enabling edge-case testing without real-world risk.
CLIP vision encoders vulnerable to typographic attacks; training-free concept localization proposed for robustness in safety-critical applications.
G-RRM neuro-symbolic framework guides symbolic solvers with symbol-equivariant recurrent reasoning models for constraint satisfaction problems.
VRRL reinforcement learning framework improves vision-language model self-reflection by grounding corrections in visual inputs for out-of-distribution robustness.
EADP framework addresses token pruning in VLMs via entropy-aware structured compression, reducing inference cost while preserving dense instruction comprehension.
Audio feature analysis from LibriVox correlates narrator acoustics with audiobook consumption; limited direct AI relevance.
TestEvo-Bench executable benchmark evaluates test-code co-evolution for automated testing agents; mined from real repositories.
Polymarket study finds human-AI forecasting outcomes driven by human capital, not model performance; majority of users either defer or rubber-stamp.
Task-Agnostic Pretraining decouples motor learning from semantic alignment in Vision-Language-Action models to reduce expert demonstration bottleneck.
Scaling laws investigation suggests LLM social simulation fidelity may not improve with compute scaling; proposes dedicated research attention.
OrbitQuant data-agnostic quantization for diffusion transformers; uses randomized permuted block-Hadamard rotation for stable timestep-invariant inference.
Neuron-aware data selection improves LLM self-distillation in annotation-free post-training without domain-specific labels.
Philosophical critique of LLMs as cultural measurement tools, arguing models constitute rather than neutrally measure cultural phenomena.
Theoretical analysis of distributed self-supervised learning robustness under non-IID data; Masked Image Modeling shows inherent robustness.
Quantum algorithm study on stabilizer state testing with limited coherent memory; theoretical contribution tangential to LLM frontier.
EvoPolicyGym benchmark for autonomous policy evolution; evaluates how agents iteratively improve executable policies under fixed budgets.
Adaptive Transformer for time series forecasting with extreme-event handling; specialized application to hydrologic streamflow prediction.
Observational study: reasoning effort, not tool access, drives first-try reliability in agentic code generation across model generations.
Geoffrey Litt argues developers must maintain deep code comprehension to effectively collaborate with AI coding agents and avoid cognitive debt.
Comparative evaluation of frontier LLMs (GPT, Claude Opus, Gemini, GLM) for automated Linux/bash exam grading using cognitive taxonomy.
WorldSample: data augmentation framework closes real-synthetic loop for robot RL via world models to reduce costly physical interactions.
Hybrid quantum-classical federated learning for multi-agent activity recognition; quantum enhancement for heterogeneous sensor fusion.
NeuFS uses neuron-level signals to identify high-value samples for few-shot LLM adaptation, reducing annotation costs in domain specialization.
LIME generates language-conditioned camera motion for autonomous robots, treating egocentric viewpoint control as a first-class action.
Bibliometric study shows NLP research authors migrating from ACL to ML venues post-LLM era, with 19.2pp decline in established author participation.