smevals - a small eval suite for evaluating models, prompts, and harnesses
Simon Willison releases smevals, an open eval framework for benchmarking models, prompts, and inference harnesses across configurations.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
Simon Willison releases smevals, an open eval framework for benchmarking models, prompts, and inference harnesses across configurations.
India's app market generated a record $345 million in Q2.
Had the hacks used conventional methods, someone would likely go to prison.
“What on earth is Google doing?” Misinformation fears spur walk-back of AI tool.
Slack Emoji Maker tool: AI-powered 128x128 transparent image editor for creating custom Slack emojis.
A tool that allowed anyone to generate fake AI-generated imagery and superimpose it over real Google Earth maps quickly spurred backlash.
Pompeii reimagined with Nano Banana 2 in Google Earth. | Image: Google Google has shut down Google Earth feature it launched Thursday that allowed users to edit satellite images with text prompts using AI. The tool essentially let users create AI deepfakes of the real world using text prompts; Digital Digging's Henk van Ess, for example, intentionally generated images adding things like refugees near the Mexican border and a bomb crater by a hospital in Gaza. On Thursday, Google's initial response to van Ess noted that images generated with Nano Banana 2 in Google Earth included a digital wat...
LemonLime’s CEO got “carried away” with tattoo gimmick.
Gaps in laws may help Pennsylvania high school escape AI nudes scandal.
TokTier eliminates redundant tokenization in agentic LLM serving by caching token state, reducing tokenization overhead from 64% to <5% on agent workloads.
ExtractBench introduces first benchmark for schema-guided document extraction with 4,869 pages across 370 enterprise documents, measuring accuracy, completeness, grounding, and cost.
DP-GRAMS recovers density modes under differential privacy constraints via noisy mean-shift on Hölder-class distributions, advancing theory without direct AI systems application.
SignMuon applies sign compression to Muon optimizer but exhibits theoretical convergence failures; explores limits of error feedback in matrix-aware optimization.
Freeze-then-select method decouples PDE discovery from neural optimization via structured field adapters and Stability-Validated Weak Selection for sparse observations.
GQ-FSL combines stochastic quantization with federated split learning to reduce energy consumption in edge DNN deployment on resource-constrained mobile devices.
FDD-ON ontology standardizes fault detection/diagnostics for HVAC systems, addressing interoperability and domain knowledge representation for building maintenance applications.
Iterated learning experiments show language compositionality emerges from frequency-structured meaning spaces under transmission bottlenecks.
Now a text prompt is all it takes to generate reality-warping images using Google Earth's satellite, aerial, and 3D imagery, like these images generated by Digital Digging's Henk van Ess that show "refugees near the Mexican border" and a bomb crater near a hospital in Gaza. Google responded to Digital Digging's AI-altered images, saying, "We take misinformation seriously - every image created with Nano Banana in Google Earth includes the SynthID digital watermark, so if someone is unsure about an image, they can ask the Gemini app or use Lens in Search to see if the image was AI-generated." I...
AgentHPOBench evaluates LLM agents on sequential hyperparameter optimization across 30 ML tasks, assessing experimental interpretation and adaptive decisions.
Proposes Socratic Test framework for automated, multimodal conversational assessment integrating dynamic assessment and Bloom's taxonomy.
CENDRe method extracts temporal and spectral concepts from CNN time-series models via joint frequency-domain analysis and improved localization.
Analyzes when on-policy interaction and value functions improve imitation learning over behavior cloning under representational constraints.
Introduces EPC score metric for evaluating XAI explanation fidelity by balancing feature sparsity and model performance preservation.
Snapchat has adjusted it recommendation systems to ensure that only videos created by real people are eligible for Spotlight recommendations, taking a stance against AI slop.
WCM world critic model for Vision-Language-Action RL uses observation history supervision to improve robotic manipulation post-training.
QASP predicts per-query recall curves via supervised regression to derive adaptive vector search policies without iterative invocations.
Several record labels, including the big three - Universal Music Group, Sony Music, and Warner Music Group - have proposed rules regarding chart eligibility for AI songs. In short, they wouldn't be. The proposal goes quite a bit further than a labeling proposal put forth by the RIAA, the International Federation of the Phonographic Industry (IFPI), SAG-AFTRA, and others. That would create a set of standardized labels for AI-generated and AI-assisted music. The labels' proposal would require songs be clearly labeled, but it would also keep them off international charts unless they met specific...
FriendBench evaluates 26 multimodal LLMs on dyadic familiarity inference from video; top models match human performance but via different mechanisms.
Proposes optimization-path organization framework for parameter-efficient LLM fine-tuning reducing catastrophic forgetting across heterogeneous task sequences.
Policy gradient convergence analysis for multi-armed bandits in continuous-time RL with diffusion environments; regret bounds derived.