OpenForgeRL: Train Harness-native Agents in Any Environment
OpenForgeRL enables end-to-end training of harness-native agents with open infrastructure, addressing limitation of complex inference harnesses like Claude Code.
Every story tagged with this topic, ordered by date.
OpenForgeRL enables end-to-end training of harness-native agents with open infrastructure, addressing limitation of complex inference harnesses like Claude Code.
DONDO releases 26 open w2v-BERT speech recognition models for African languages spanning six countries, trained on religious text corpora.
Open-source evaluation framework for open-weight LLM agents on longitudinal data tasks, addressing privacy constraints in research deployments.
Laguna S 2.1, a 118B MoE model from Poolside AI, achieves Deepseek v4 Pro performance at lower cost than v4 Flash.
Poolside AI co-CEO Eiso Kant describes building a model factory enabling efficient training of 118B MoE models competitive with 1T open-weight alternatives.
Security researcher Thomas Ptacek claims open-weights 2025 models could execute sandbox escapes and network reconnaissance without frontier capabilities.
18 medical image encoders on 650k radiographs show self-supervision drives representational convergence more than clinical labels.
Orchestrated open-weight small LLMs achieve malware analysis performance competitive with frontier closed-weight models at lower computational cost.
CircuitKIT open-source library unifies circuit discovery, evaluation, and intervention workflows for mechanistic interpretability with automated contrastive prompts.
Ben Thompson proposes US law to legalize model distillation and data collection as fair use, addressing licensing hypocrisy and competitiveness vs. Chinese models.
Stratechery argues U.S. frontier labs face minimal threat from Chinese models; policy should prioritize open-weight domestic alternatives instead.
Sam Altman email (Oct 2022) reveals OpenAI planned GPT-3-class open-weight model for consumer hardware to preempt Stability AI.
Study evaluates open-weight LLMs for extracting structured CVE threat data from autonomous vehicle vulnerability text.
Kimi K3 2.8T-A50B released as largest open-weight model with Opus 4.8-class performance at Sonnet 5 pricing.
Moonshot AI releases Kimi K3 (2.8T params), claims top performance vs. Claude Opus 4.8 Max and GPT-5.5, promises open-weight release by July 2026.
Thinking Machines Lab releases Inkling, a 975B-parameter open-weights MoE multimodal model trained on 45T tokens.
Thinky releases Inkling, a 975B multimodal open-weights model under Apache 2.0, with a smaller 276B variant.
Pythia multi-agent system for autonomous clinical symptom extraction using open-weights LLMs without fine-tuning.
Open-weight reasoning models fine-tuned via RLVR for thermal energy storage control, achieving building-scale load shifting with 30 prompts.
Cohere releases Tiny Aya Expedition, a multilingual model supporting 70+ languages for on-device and educational AI applications.
Soofi S 30B-A3B: open-source MoE-Mamba hybrid for German/English with 3B active parameters, matches 14-27B dense models on benchmarks.
LLM4SDM evaluates open-source smaller models on clinical decision-making assessment, comparing privacy-preserving local deployment vs. commercial models.
Cohere releases open-source Arabic speech recognition model for enterprise transcription across Arabic dialect variants.
Tencent releases Hy3, a 295B-param MoE model with 21B active params under Apache 2.0, claiming performance parity with 2-5x larger open-source competitors.
Linear probes decode remaining output length from LLM hidden states across 7-8B open-weight models, revealing internal response-length estimation.
Current AI launches Gap Map v0.1, an index of 421 open-source AI products across models, tools, datasets, and hardware, backed by $400M committed capital.
Simon Willison's June 2026 newsletter covers Claude Fable 5, GPT-5.6, GLM-5.2 open weights, and US export restrictions.
HaloGuard 1.0 releases open-weights constitutional safety classifier achieving state-of-the-art multilingual prompt-safety performance at 1/10 model size.
MultiSynt/MT releases 4.8 trillion tokens of open synthetic parallel pre-training data across 36 European languages via Tower+ and OPUS-MT translation.
AI Engineer World's Fair coverage: agent loops, software factories, forward-deployed engineering, and open model adoption emerging as key themes.
Quantized open-weight LMs on 16GB VRAM match proprietary APIs on database tasks at lower cost and latency.
Google UK publishes economic impact report on AI adoption and productivity benefits, positioning open-weights models as tools for broader UK workforce enablement.
DeepReinforce releases Ornith-1.0, MIT-licensed open-weights model (9B–397B variants) for agentic coding, built on Gemma 4 and Qwen 3.5, achieving SOTA on coding benchmarks.
Multi-agent system using open-source LLMs for automated disinformation detection and fact-checking at scale across social media.
Paris 2.0: first decentralized video generation model trained without GPU clusters, extending prior Paris 1.0 image work.
User reports Qwen3.6 35B outperforms Gemma4, GLM 4.7 Flash, others for local agentic tasks; seeks comparable MoE alternatives.
MiniCPM5-1B released on HuggingFace: 1B-parameter model from CPM team, likely competitive efficiency benchmark for edge deployment.
Financial Times reports Heretic tool removes guardrails from Meta's Llama 3.3 in <10 minutes; 3,500+ decensored variants downloaded 13M times.
Numind releases NuExtract3, open-weight 4B multimodal VLM for document extraction and Markdown conversion under Apache-2.0.
MiMo-V2.5-coder released as open-weights coding model alternative to Qwen and DeepSeek for 128GB+ systems.
Elon Musk announces 0.5T parameter Grok model planned for next year, with open-weights release.
hipEngine: open-source ROCm-native inference engine for Qwen 3.6 MoE on AMD RDNA3 GPUs (7900 XTX, Strix Halo).
BitCPM-CANN demonstrates 1.58-bit ternary quantization training on Huawei Ascend NPUs, addressing extreme low-bit LLM deployment outside CUDA.
Reddit discussion comparing inference speed/quality tradeoffs between Qwen3.6-35B and Gemma4-26B on consumer GPU hardware.
Community finetune of Qwen 3.6 35B with quantized weights; testing on consumer hardware shows stability at 200k context.
Community-built open-source TTS benchmark suite with Windows/Mac results; Linux results pending, covers known local TTS tools as of May 2026.
Reddit discussion questioning utility of uncensored models for RAG applications; user reports stability issues vs. base models.
llama.cpp server adds native tool support (shell execution, file ops) via experimental --tools flag.
Chrome extension enables local inference of Gemini Nano (Gemma) on CPU-only systems, ~20 tokens/sec on laptop.
Developer refactored 120-file FastAPI service using DeepSeek V4 and Hunyuan with 80x cost savings vs Opus; open-weight models matched Opus latency but introduced production bugs.