Inducing Task Models from Computer-Use Traces
Method induces reusable symbolic task models from passively recorded computer-use traces (screenshots, mouse/keyboard actions) for agent learning and organizational audit.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
Method induces reusable symbolic task models from passively recorded computer-use traces (screenshots, mouse/keyboard actions) for agent learning and organizational audit.
AI4AI-Bench isolates LLM agent capability to design training algorithms for recursive self-improvement, beyond hyperparameter tuning or data collection.
Pandora's Box framework optimizes routing across heterogeneous AI systems, balancing cheap noisy value estimators against expensive accurate ones.
BERT-LER pretrained on 75M-patient EHR dataset uses discrete laboratory tokens with percentile binning and Integrated Gradients for interpretable clinical prediction.
MidTool: open pipeline synthesizing tool-use training data from APIs, web, and code to improve agentic LLM capabilities via mid-training.
Statistical inference method for dictionary learning under calibration uncertainty; niche signal-processing contribution.
Affected users told TechCrunch they were using Grok Lite, and noticed the issues as early as Wednesday morning.
Study identifies seven measurement artifacts in LLM self-improvement claims; Qwen3-8B LoRA self-training masks false gains without proper null controls.
Causal graph learning from sleep apnea test data reveals population-stratified breathing dynamics; medical domain application.
ChatGPT and other AI models are now authoring and editing much of the new web.
IAR: three-stage post-training framework (inject, align, recover) for internalizing bounded document corpora into LLM parametric knowledge.
Systematic benchmark of LLM semantic cache eviction policies (FIFO, LRU, LFU, ARC) shows LFU dominates across settings with <1% margin.
Study compares task-level vs. subtask-level skill induction in LLM agents; finds subtask-level code skills transfer more reliably across domains.
ML classifier for early detection of Solana memecoin rug pulls using liquidity and social signals; blockchain fraud prevention.
DICS: clustering-based approach to reduce candidate split search in decision tree training via data-informed centroid selection.
LLM learns adaptive reasoning effort allocation (NoThink/Short/Long) via GRPO to optimize test-time compute per problem difficulty.
Transfer learning framework for nonparametric regression using deep ReLU networks with group-specific offsets and L2 error bounds.
Ramp has launched its own AI model routing service, dubbed Router, that lets users and companies use and switch between various large language models via an API.
QUASAR: hybrid quantum-classical neural network for SAR satellite physical-layer authentication using RF fingerprinting at X-band frequencies.
RuleMaze benchmark evaluates MLLMs on rule-compliant spatial planning in mazes with natural-language constraints.
Economic model of post-AGI corporate ecosystems with AI/robotic producers and consumers achieving demand closure independent of human consumption.
Prompt-conditioned channel attention mechanism for hierarchical feature modulation in medical image segmentation tasks.
Machine-learning surrogate waveforms accelerate gravitational-wave parameter estimation for eccentric orbits and high-mass-ratio binaries.
InsufficiencyBench evaluates LLM legal advice recognition of materially missing information in underspecified queries across eight failure categories.
ML system classifies Electronic Navigational Chart changes as critical or non-critical risks to reduce manual maritime safety review labor.
Daedalus-150M: 150M-parameter convolution-attention hybrid optimized for 4-bit CPU inference with fixed-width convolution cache, trained on 59.9B tokens.
Meta is bringing Pocket, its experimental AI-powered app for creating and sharing interactive games, to users across the U.S. after quietly testing it in Brazil.
ContractScrub benchmark evaluates LLM performance on legal contract final review for errors, inconsistencies, and NER across long documents.
MemTrapBench evaluates how retrieved memories distort LLM reasoning even when faithfully stored, identifying memory-induced cognitive failure modes.