RARE: Decoupling Representation Steering from Expert Routing in Mixture-of-Experts Language Models
RARE framework enables representation engineering steering for MoE LLMs by decoupling routing from behavior-directed latent modifications.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
RARE framework enables representation engineering steering for MoE LLMs by decoupling routing from behavior-directed latent modifications.
Master-level lecture notes on numerical linear algebra foundations for PDEs, ML, and data assimilation.
GPU-optimized genetic algorithm for TSP with irregular memory access patterns; domain-specific optimization unrelated to frontier AI.
Agent memory poisoning via false statements reduces accuracy 0.85→0.30; content screening fails to detect 360 poisoned memories.
Ontology framework for tracking and managing AI models/datasets across organizations; metadata and governance infrastructure.
Zeroth-order optimization for spiking neural networks on in-memory computing accelerators; hardware-specific fine-tuning technique.
LLM augmentation with semi-structured data for predictive political QA; actor stance and vote prediction from external resources.
P3Bench: benchmark for personalized privacy control in LLMs via attention head intervention; user-specific disclosure preferences.
Cross-agent specification portability: Claude, Gemini, Copilot on Oracle-to-PostgreSQL migration; 380/1006 successful executions.
Curriculum-aware time-series imputation for clinical physiological data; addresses realistic gap patterns in medical records.
PUN protocol: plausible unknown names for LLM evaluation; controls for memorization vs. factuality in person-name benchmarks.
Q-Planning: off-policy Q-function enables self-improving robot policies beyond behavior cloning; scales to multi-billion-parameter visuomotor models.
SENTRY applies XGBoost and RAG to replace questionnaires in IT change risk assessment for financial institutions.
Tydra hybrid Transformer-SSM architecture reduces TabPFN inference time 30% on tabular data while maintaining accuracy.
Matt Webb describes using ChatGPT as interactive tutor to learn quaternions for AR app development, arguing AI augments rather than replaces learning.
AI factories are power-constrained industrial systems. The question is no longer how many GPUs fit in a data center, but how much AI output each available... AI factories are power-constrained industrial systems. The question is no longer how many GPUs fit in a data center, but how much AI output each available megawatt can deliver. For AI inference workloads, this makes application-level performance per watt the key metric for measuring AI factory efficiency. Not every megawatt translates to revenue-generating compute. Power distribution, cooling… Source
OCT imaging dataset for quantifying cochlear fibrosis in hybrid cochlear implants to improve hearing outcomes.
ARDLS method stabilizes Analytic Hierarchy Process priority vector derivation under high inconsistency via regularization.
NSPIN neurosymbolic framework combines LLMs and symbolic methods to induce probabilistic planning models from surgical narratives.
DAMOS learns speech quality assessment by explicitly localizing distortion regions rather than utterance-level MOS labels.
SRL-MPC combines RL with MPC for shape-aware navigation of heterogeneous robots in dense crowds without geometry simplification.
OAttention replaces binary attention masks with token presence coefficients to enable graduated token participation in Transformer layers.
Thermo-FL adds thermal-aware control and Byzantine robustness to federated LoRA fine-tuning of LLMs on edge devices.
SPaRC benchmark shows visual scaffolds significantly improve VLM spatial reasoning on grid-based planning tasks.
Human-JEPA extends JEPA architecture to video for dense human perception and future motion prediction without collapse.
AID-Guard introduces stateful authorization protocol for tool-using agents to prevent duplicate effects and state conflicts.
HIERA proposes hierarchical search-space planning for automated GPU kernel optimization across implementation variants.
Graph engineering framework enabling LLM agents to coordinate heterogeneous expertise and parallel subtasks via system-level orchestration.
Winder applies phase-equivariant self-supervised learning to cardiac ECG representation via closed-form transport operators.
COEC proposes alternating orthogonal compensation for training-free structured pruning of LLMs to reduce inference cost.