Generative Skill Composition for LLM Agents
Generative Skill Composition uses LLM-guided retrieval to select and compose skills for complex agent tasks without exposing full skill library.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
Generative Skill Composition uses LLM-guided retrieval to select and compose skills for complex agent tasks without exposing full skill library.
Startup Acti is betting the smartphone keyboard is the next home for AI assistants. Its new keyboard for iOS and Android works across apps and lets users create custom AI-powered shortcuts using natural language.
FLORA applies deep learning to LiDAR data for forest attribute prediction under heterogeneous acquisition conditions.
Semantic Reference Frames (SemRF) standardize residual-stream analysis across LM layers to distinguish computation from measurement drift.
Automated Background Swapping reduces classifier reliance on spurious background correlations through synthetic data augmentation.
TRIAGE framework improves credit assignment in agentic RL by classifying action segments into semantic roles, addressing limitations of uniform outcome-based reward.
FedLAB enables federated learning of multimodal graph foundation models with semantic traceability while preserving privacy across decentralized clients.
Skill distillation approach scales browser agent training via behavior cloning from implicit priors in human browsing traces rather than low-level operation.
CoMet decomposes uncertainty sources in multimodal LLMs via context and multiplicity factorization for better calibration in open-ended tasks.
Surrogate Fidelity evaluates when mechanistic interpretability findings from open models transfer to closed model predictions, attributions, and representations.
AxDafny framework uses verifier-guided repair for agentic code generation with proofs in Dafny; introduces LiveCodeBench-Pro-Dafny benchmark of 250 problems.
Simon Willison reviews an interactive quiz that classifies users into 30 AI ethics archetypes based on 29 questions.
Random Reshuffling provably dominates SGD in optimization theory, formalizing why this practical heuristic outperforms classical stochastic gradient descent.
PolicyGuard converts organizational policies into executable neuro-symbolic compliance review engines with interpretable logic rules and LLM integration.
GPU-accelerated query engines are often constrained by memory and I/O bandwidth. NVIDIA hardware advances—including high bandwidth memory (HBM), NVIDIA... GPU-accelerated query engines are often constrained by memory and I/O bandwidth. NVIDIA hardware advances—including high bandwidth memory (HBM), NVIDIA NVLink-C2C, and dedicated decompression engines featured in NVIDIA GB200 NVL4—help remove these bottlenecks by increasing effective storage capacity, accelerating data movement between CPUs and GPUs, and speeding data access without consuming… Source
Self-generated QA for model distillation exhibits fragility in coverage and evidence selection, showing implicit policy biases degrade downstream fine-tuning.
Radial suppression of activation-space inflation during training accelerates generalization on algorithmic tasks by inducing data-dependent weight regularization.
MADreMIA framework amplifies membership inference attacks on generative models via chained regeneration, enabling scalable privacy auditing without shadow models.
Comparative study of genetic programming population initialization methods for symbolic regression finds no significant differences across techniques on synthetic benchmarks.
GR2 demonstrates LLM-based re-ranking for industrial recommendation systems, addressing final-stage funnel optimization typically underexplored in LLM recommendation work.
LUNA proposes LBS-free neural animation model using transformer-based motion regression and 3D Gaussian deformations for photorealistic human avatars from monocular images.
A new teaser trailer confirmed that Wonka's The Golden Ticket will premiere on Netflix on September 23rd, following its Squid Game reality show in the trend of creating real competitions based on fictional torture scenarios. While the sets seen in the trailer are real and not some Glasgow-style AI fakes, the voiceover is AI-generated. Deadline reports that Netflix worked with AI audio company ElevenLabs with consent from Wilder's family, after working on productions re-creating the voice of Michael Caine and Stan Lee. It also continues Netflix's 2021 partnership with the Roald Dahl company, a...
DigitalCoach dataset of 72 expert-novice coaching sessions reveals models coach differently than humans—more direct instructions, fewer explanations—raising generalization concerns.
TreeAgent multi-agent system combines expert decision trees with Vision-Language Models for tree height bias classification in forestry remote sensing with voting-based stochasticity mitigation.
Relay-assisted semantic communication systems leak privacy: intermediate relays can infer semantic meaning and reconstruct signals without source data access.
MECoBench evaluates multimodal LLM collaboration in embodied environments, finding collaboration improves task completion but benefits depend on balancing gains vs. overhead.
Signed-permutation coordinate transport framework enables steering vector and SAE alignment for RMSNorm transformers via sign-marginalized Hungarian algorithm.
Anthropic's Claude Science is a workbench that gives scientists one environment to do computational research, saving them from the need to bounce between databases, pipelines, and tools.
A year in, National Design Studio delays plan to update government web standards.
shot-scraper 1.10 adds video recording capability via storyboard.yml to help coding agents demonstrate their work with Playwright automation.