Sophistication in GenAI Use: Field Evidence from a Large Firm
Field study of 713k prompts from 4k employees reveals senior staff use GenAI more sophisticatedly; variance across functions.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
Field study of 713k prompts from 4k employees reveals senior staff use GenAI more sophisticatedly; variance across functions.
Anthropic launches free Claude access program for 10,000 verified scientists with subsidized team subscriptions.
Speaker anonymization suppresses speaker-identifying attributes from speech while preserving linguistic content and quality. We propose repurposing XTTSv2, a multilingual voice cloning model trained on 27k hours of speech, for speaker anonymization without retraining. Our key insight is that XTTSv2's voice cloning capabilities preserve prosodic structure independently of speaker identity, enabling voice conversion by conditioning on a pseudo-speaker. We introduce an iterative refinement strategy that balances privacy and utility by maximizing a harmonic mean of speaker dissimilarity and intel...
RCMN framework operationalizes misleadingness in public discourse across five dimensions including intent and reader interpretation.
Evolution Strategies show broader reasoning coverage than GRPO in LLM post-training with lower memory overhead.
INTENT-AS-A-TOOL monitors agentic misalignment by adding tools for explicit intent signaling during agent reasoning.
PAWBench evaluates whether video generation models capture correct distributions of physical outcomes, not just plausible frames.
Workflow for recovering 18th-century essay republication from fragmented text-reuse evidence across ECCO and newspapers.
Eval-aware CoT framings predict model compliance differently; capabilities-framing on Qwen3-32B strongly predicts rule-following.
Information floor metric separates model gap from missing context in block drafting; tests four open-weights and frontier API models.
BTS-AgentBench pipelines industrial telemetry logs into deterministic multi-turn agent tasks with 532 episodes.
Finite-sample convergence proof for quantile temporal-difference learning in distributional RL via M-matrix analysis.
On Nvidia's earnings call Wednesday, CEO Jensen Huang casually announced the company had "achieved AGI," one of the tech industry's ultimate goals some of its biggest players have spent years chasing. Almost immediately, Huang dismissed the coveted milestone as "senseless." He's right. For the supposed finish line of the AI race, there is no consensus on what artificial general intelligence means, let alone how we'll know when we've actually got there, which makes achieving it equally arbitrary. Asked about OpenAI's pursuit of AGI, Huang said that when it comes to Nvidia, "for many tasks, we ...
HarnessLens automates agent harness evolution via behavior-aware verification, reducing redundant rollouts across task space.
Google DeepMind releases Gemini 1.1 Flash with enhanced control features for developers building applications.
Pre-registered study reveals that difference-in-differences designs on bounded rating scales can spuriously manufacture bias effects in LLM-judge audits via differential censoring.
QuantumBoostNet hybrid classical-quantum model for cardiac ultrasound view classification in medical imaging.
LLMs can synthesize near-optimal algorithms for operations research problems (inventory control, queueing, assortment) given problem class descriptions and parameter ranges.
The updates indicate that Google is looking position to AI Mode as an AI travel agent of sorts, as it's moving beyond simply helping users find information to actually handling parts of the trip-planning and booking process itself.
Google Search adds AI-powered travel booking features: hotel reservations, airfare tracking, and loyalty rewards integration.
Proposes expert critic-sourced network adjacency as music recommendation signal beyond collaborative filtering and acoustic content.
MM-Spectrum: sparse Mixture-of-Experts framework for multimodal molecular structure elucidation from heterogeneous spectroscopic signals.
TADP: task-aware deformable prediction method for single-stage 3D object detection with multi-scale feature aggregation.
BrailleBench: new benchmark evaluating LLM comprehension of Braille across five criteria with 5,570 instances for accessibility assessment.
Naive Prompt Optimization (NPO): lightweight single-lineage prompt optimization via teacher rollouts, challenging need for complex prompt search methods.
SCIT (Suffix Cache Interchange Test): causal protocol for testing which transformer components carry latent reasoning in chain-of-thought models.
Two-level framework analyzing agentic data generation requirements: consistency across environments, tasks, interactions and quality vs. quantity tradeoffs.
Study on architectural choices for latent-space transition prediction in robot World Action Models, comparing Transformer-based approaches to operator-structured designs.
Neural operator architecture improvements for enforcing Dirichlet boundary conditions in PDE solution approximation.
Circuit Condensation post-training method that prunes model causal circuits to isolate behaviors, enabling mechanistic interpretability inspection.