MemSyco-Bench: Benchmarking Sycophancy in Agent Memory
MemSyco-Bench benchmarks agent sycophancy in memory retrieval, measuring how retrieved memories bias reasoning and decision-making.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
MemSyco-Bench benchmarks agent sycophancy in memory retrieval, measuring how retrieved memories bias reasoning and decision-making.
GSRQ applies gain-shape residual quantization to push KV cache to sub-1-bit regime, addressing centroid shrinkage in high-dimensional codebook learning.
Multi-agent LLM framework automatically generates and verifies reaction rules across 665k patent reactions for chemistry synthesis planning.
Theoretical characterization of separable graphical models with mixed edge types for encoding independence structures.
LLM agent collectives with persistent memory and tools exhibit emergent complexity while remaining interpretable substrates.
Test-time control framework DART-VLN mitigates memory decay and loop inefficiencies in vision-language navigation agents.
EchoRisk: multicentre longitudinal echocardiography dataset with cardiotoxicity labels for automated breast cancer treatment risk stratification.
Study shows moderate LLM agent personality expression outperforms extremes on trust and goal-adoption in conversational behavior-change tasks.
Post-hoc SFT training makes model-organism interpretability unrealistically easy; 54-model suite shows methodology strongly affects white-box evaluation validity.
RF drone identification benchmarks leak data via cross-validation on continuous recordings; theory and measurement expose inflated reported accuracies.
FinKG-News framework constructs news-anchored financial knowledge graphs for explainable credit risk report generation via in-context learning.
SEAHORSE: unified benchmark framework standardizing preprocessing, splits, and evaluation for spatiotemporal point process neural models.
PedNStream: open-source Python simulator for macroscopic pedestrian network flow based on Link Transmission Model with stochastic dynamics.
Graph-based, training-free framework for reading order inference in complex historical documents using lightweight language-model signals.
Generative model-based particle filtering for Bayesian data assimilation in high-dimensional state estimation problems.
Mathematical framework extending Cover's function-counting theory to characterize binary classification on low-dimensional data structures.
Survey chapter on LLM mechanisms: Transformer architecture, attention, emerging capabilities, and layer-wise implementation of reasoning.
Evan Feinberg and Sergey Edunov discuss diffusion models for drug discovery at Genesis Molecular AI, including PEARL's OpenBind performance and protein co-folding advances.
Logit-Contribution Scoring method identifies attention heads performing non-literal synthesis in long-context LLM retrieval tasks.
Benchmark of foundation models vs. radiomics for lung CT analysis, isolating contributions of feature extractors, classifiers, and segmentation.
KnowledgeDebugger: GUI tool for knowledge localization and editing in Transformers, integrating EasyEdit library with no-code interface.
Let’s start with a game. Open up your chatbot of choice—Claude, ChatGPT, Gemini—and type “Give me a random number between 1 and 10.” You’re going to get 7. Almost always. Now type “Another” and you’ll get 3 or 4. Type “Another” again and you’ll get 8 or 9. That won’t work every time—but if it…
Multitask learning framework for mixed-type outcomes using shared sparsity and monotone transformation-invariant loss functions.
Warp CEO Zach Lloyd argues automated software factories will become standard for major projects, outlining preparation strategies for engineers.
SWE-Doctor: LLM-based code agent using multi-faceted bug reproduction tests for runtime diagnosis to improve software patch generation.
LLM-powered agent simulates semantic human trajectories in zoned environments to reduce costly real-world data collection.
Venice AI is already profitable, with annualized run-rate revenues of over $70 million, CEO Erik Voorhees said.
ML pipeline detects stress from speech using speaker diarization on Trier Social Stress Test data.
Google's 24/7 agentic assistant, Gemini Spark, comes to Mac alongside other improvements, like real-time tracking and support for more apps.
TRCGL-Net addresses long-tail chest X-ray classification via diffusion-based augmentation and label co-occurrence modeling.