AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing
AuK, open-source multimodal model for unified speech generation and editing via natural-language instructions and audio context.
Every story tagged with this topic, ordered by date.
AuK, open-source multimodal model for unified speech generation and editing via natural-language instructions and audio context.
Study characterizes interaction between inference-time activation steering and weight-only quantization (INT8, NF4) on 7-9B open-weight LLMs.
Study of 3,471 uncensored open-weight models on HuggingFace (Jan 2024–Mar 2026), tracking safety guardrail removals and redistribution persistence.
Causal taxonomy distinguishes deceptive behavior from deceptive mechanisms in language models, tested on open-weight families.
Open-weight transformers show logical validity representations remain decodable from hidden states despite near-chance behavioral performance.
Google announces August 2026 AI updates; article lacks specific details on Gemma releases or capabilities.
Causal analysis across 9 open-weight models shows quantization damage is distributed globally, not localized to task circuits or weight statistics.
HiveTraceGuard-Pro, a 0.6B LoRA guardrail model tuned on Qwen3, detects Russian and English prompt injection and jailbreaks with binary safety scoring.
Tencent releases Hy4, a 770B open-weight LLM with 49B active params and 1M token context, 2.6× larger than Hy3.
Open-source tendon-driven robotic hand simulator for dexterous manipulation learning with underactuated transmission dynamics.
Three-stage post-training recipe (acquire, repair, preserve) for 2B open-weight dialogue game agents using diagnostic error analysis and RL.
Linear probe analysis of how open-weight LLMs organize Moral Foundations Theory categories in representation space.
Puro-2B: open-source 1.5B LM trained on RTX 5090 for <$5k, targeting cost-accessible pretraining.
Qwen releases Qwen3.8-Flash-Next, a 125B-parameter MoE model with 6B active tokens and multimodal capabilities, previewing Qwen4 architecture.
Z.ai CEO Jie Tang discusses GLM 5.3 model and post-training scaling laws as alternative to parameter growth.
Glean CEO explains model routing as cost-control mechanism driven by frontier model pricing and open-weights adoption, improved via scaled human feedback.
Mojo programming language released as open source under Apache 2 license after 1.0 stable release, pivoting from Python superset goal.
Google announces July 2026 AI updates; article lacks specific details on models, features, or benchmarks.
Alibaba releases Qwen 3.8 Max (2.4T params) and 27B open-weight models optimized for coding and collaboration tasks.
Analysis of 18 open-source LLMs showing cultural bias in mythology knowledge; models encode cross-cultural distinctions in residual streams but fail to decode non-Western traditions.
Antares: compact LLMs (350M–3B) for agentic vulnerability localization via SFT and RL on cybersecurity reasoning over code.
Microsoft-led open letter signed by 235 AI companies including NVIDIA and OpenAI argues against US government restrictions on open-weight models on safety grounds.
Gaokerena: compact Persian-language medical LLM family trained on 90M-token corpus for low-resource healthcare deployment.
Tevatron 3.0 integrates Megatron-Core for efficient MoE reranker training, enabling billion-scale cross-encoder + distillation workflows on academic budgets.
DeepSeek releases V4-Flash-0731, a 304B parameter model with enhanced agentic capabilities, outperforming larger competitors at $0.14/$0.27 per million tokens.
Podcast discussion on open-weight model competitiveness, cybersecurity risks, and AI leadership policy letters signed by industry leaders.
Simon Willison releases smevals, an open eval framework for benchmarking models, prompts, and inference harnesses across configurations.
DenseOn and LateOn: open-source 149M-parameter retrieval models trained on 1.88M supervised pairs; competitive on multilingual and code search.
Method to detect CSAM-generating LoRAs from weight fingerprints (singular vectors) without generating outputs, enabling safer moderation.
Kimi K3 open-weights model released amid broader industry discussion on open model availability and strategy.
Moonshot releases Kimi K3 weights (2.8T params, 1.56TB) with modified MIT license requiring attribution for products >100M MAU.
Anthropic publishes official stance on open-weights model releases, addressing trade-offs between transparency, safety, and competitive positioning.
Causal-TS: open-source Python library for causal discovery in high-dimensional nonstationary time series with GPU-accelerated conditional independence testing.
Kimi K3: 2.8T MoE model with 104B active params, 1M context window, Delta Attention, 2.5x scaling efficiency over K2.
ELMOD: 2.7B German-language model optimized for mobile deployment using public data and morphology-aware preprocessing.
Byte-Prefix Marginalization method for cross-tokenizer on-policy distillation of open-weight LLMs with incompatible vocabularies.
OpenForgeRL enables end-to-end training of harness-native agents with open infrastructure, addressing limitation of complex inference harnesses like Claude Code.
DONDO releases 26 open w2v-BERT speech recognition models for African languages spanning six countries, trained on religious text corpora.
Open-source evaluation framework for open-weight LLM agents on longitudinal data tasks, addressing privacy constraints in research deployments.
Laguna S 2.1, a 118B MoE model from Poolside AI, achieves Deepseek v4 Pro performance at lower cost than v4 Flash.
Poolside AI co-CEO Eiso Kant describes building a model factory enabling efficient training of 118B MoE models competitive with 1T open-weight alternatives.
Security researcher Thomas Ptacek claims open-weights 2025 models could execute sandbox escapes and network reconnaissance without frontier capabilities.
18 medical image encoders on 650k radiographs show self-supervision drives representational convergence more than clinical labels.
Orchestrated open-weight small LLMs achieve malware analysis performance competitive with frontier closed-weight models at lower computational cost.
CircuitKIT open-source library unifies circuit discovery, evaluation, and intervention workflows for mechanistic interpretability with automated contrastive prompts.
Ben Thompson proposes US law to legalize model distillation and data collection as fair use, addressing licensing hypocrisy and competitiveness vs. Chinese models.
Stratechery argues U.S. frontier labs face minimal threat from Chinese models; policy should prioritize open-weight domestic alternatives instead.
Sam Altman email (Oct 2022) reveals OpenAI planned GPT-3-class open-weight model for consumer hardware to preempt Stability AI.
Study evaluates open-weight LLMs for extracting structured CVE threat data from autonomous vehicle vulnerability text.
Kimi K3 2.8T-A50B released as largest open-weight model with Opus 4.8-class performance at Sonnet 5 pricing.