Grokipedia vs Wikipedia: An LLM-Based Audit of Political Neutrality along Ideologies
Large-scale political bias audit compares Grokipedia (Grok-written encyclopedia) and Wikipedia across 1,394 article pairs on neutrality.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
Large-scale political bias audit compares Grokipedia (Grok-written encyclopedia) and Wikipedia across 1,394 article pairs on neutrality.
Companies coming to market are raising money at fastest pace this century.
Diagnostic study isolates and evaluates five visual world models (DreamerV3, DIAMOND, TWISTER, Simulus, STORM) in Atari Pong.
DoorDash is opening a limited beta of dd-cli, a command-line tool that lets developers and AI agents search stores, build carts, and place orders from the terminal, marking another step toward software designed for AI agents instead of just humans.
Thinking Machines Lab releases Inkling, a 975B-parameter open-weights MoE multimodal model trained on 45T tokens.
NIFA extends FPGA-integrated ReRAM in-memory computing to support nonlinear operations for efficient ML inference.
You may have heard that OpenAI released its first piece of hardware this week. You may not have heard about the ChatGPT basketball.
Categorical framework (LINCS) addresses non-compositionality in ML via tangent category sketches and universal factorization.
Hierarchical Global Attention + tiered KV storage enables 16K-token fine-tuning on 16GB VRAM, 8× longer than dense attention baseline.
Multi-agent LLM framework using SFT+DPO enables sustained partisan behavior in political coalition simulation, circumventing RLHF neutrality bias.
AlphaWiSE: post-hoc weight interpolation maintains cross-modal alignment in CLIP during continual multimodal learning via per-tensor scalar coefficients.
Drawing on more than a decade spent helping build some of the world's most influential AI systems, including research that later informed the development of ChatGPT, Andrew Dai explains why he believes visual AI is one of the next major frontiers in artificial intelligence.
Rubrics on Trial evolves query-specific evaluation rubrics via synthetic pairwise evidence, validating rubric utility without human supervision.
Simon Willison compiled Go's mermaid-ascii library to WebAssembly for ASCII diagram rendering with color support.
Bottom-up pipeline (YOLO + LayoutReader) and Tiramisu model for hierarchical structure understanding in dense newspaper images.
Covariate balance diagnostics reveal insufficient robustness in existing offline RL studies for clinical treatment recommendation.
BrainPilot agentic research framework integrates domain expertise in neuroscience to reduce hallucination and drift in multi-step scientific reasoning.
SINDy sparse regression recovers interpretable governing equations from limited engineering data without large neural network datasets.
While everyone in AI is chasing "superintelligence," Alexandre LeBrun, CEO of Yann LeCun’s world model startup, AMI Labs, dismisses the word.
Kernel-WIS off-policy evaluator for contextual bandits combines bounded importance sampling with linearity, outperforming baselines on misspecified policies.
DriftWorld accelerates diffusion-based world models for robot planning by learning action-conditioned drift during training, eliminating expensive iterative denoising at inference.
ANet Patu-1 proposes a self-organizing consensus protocol for AI agent networks that adaptively forms coalitions, modeling network value across coordination-group sizes.
The FT reports Kimi K3 will be the largest open AI model from China, with a parameter count between 2 trillion and 3 trillion.
Parameter-efficient prompt tuning of vision models for MCI screening from drawing tests, using focal loss to handle class imbalance and data scarcity.
ArtSplit provotype explores how to quantify and attribute ownership between human and AI contributions in creative work through measurable contribution metrics.
cGAP introduces a visualization framework for high-dimensional categorical data that preserves original data matrix while enabling interpretable exploration.
Anthropic launches $50k Claude credit grants for rare genetic disease research, seeking to build AI for Science researcher community.
Empirical study on authorship calibration reveals users relying heavily on generative AI misjudge their actual authorship contribution compared to light users.
SMC-ES uses simulation-based synthesis to automatically generate formally verified control policies with performance and safety guarantees for cyber-physical systems.
LQCDMaster is a domain-specialized agent that converts natural-language lattice QCD research tasks into executable PyQUDA workflows with planning and tool use.