Daily Tech Briefing
AI 科技速览
每天 5 分钟内学习 AI。获取最新的人工智能新闻,理解其重要性,并学习如何将其应用于您的工作。
MarkTechPost · 2026/8/2 20:35:50
Thinking Machines Lab Releases Inkling-Small: A 276B Total, 12B Active Open Weights Multimodal MoE Model
AI 中文解读
AI圈又出“小钢炮”了!Thinking Machines Lab发布了一款开源模型Inkling-Small,总参数2760亿,但每次只激活120亿,个头只有自家旗舰版“Inkling”的四分之一,却在推理和编程测试上反超了“老师傅”。这相当于AI界的“团队作战”:虽然公司有2760亿名员工,但处理每项任务只派12亿精锐上场,既省算力又跑得快。
更接地气的是,它终于让“大模型不再是大厂的专属玩具”。以前跑这种级别模型得堆好几台顶配服务器,现在经过压缩优化,一张高端显卡就能搞定,成本断崖式下降。创业公司租一台云服务器就能自用,医院、银行这类对数据敏感的行业也能把模型部署在自己机房,不用担心机密外泄。
它还是“多面手”,能看图、听音频、读文档,甚至处理长达100万字的超长文本。以后语音助手开会能自动总结,客服中心能实时分析通话情绪,程序员写代码它也能当得力副手。最关键的是,模型权重完全开源免费,技术民主化的步子又迈大了一步。
Thinking Machines Lab has released Inkling-Small, an open weights Mixture-of-Experts model with 276B total parameters and 12B active. That is about a quarter the size of Inkling, which carries 975B total and 41B active parameters. The model was trained on NVIDIA GB300 NVL72 systems. It reasons natively over text, images and audio. The context window reaches 1M tokens, and thinking effort is adjustable. Weights ship under Apache 2.0 on Hugging Face.
Is it deployable
Yes, and the quantized checkpoint is why. Per the model card, the BF16 checkpoint needs at least 600 GB of aggregated VRAM. That is met by 4x NVIDIA B300 or 8x NVIDIA H200. The NVFP4 checkpoint drops that floor to 180 GB. It runs W4A4 on a single B300, which requires SM100+, or W4A16 on two H200s. Supported runtimes are SGLang, vLLM, TokenSpeed, Unsloth and Hugging Face.
That single-GPU path moves a 276B model out of frontier-lab territory. Startups can self-host on one rented B300 instance. Mid-size enterprises with existing H200 capacity can serve it without new hardware. Regulated sectors gain a private-weights option: financial services, healthcare operations, insurance, telecom and public sector. Applicable workloads include coding agents, terminal automation, and document and chart understanding. Audio widens that to call-center analytics, voice interfaces and meeting summarization.
Architecture
Inkling-Small is a 42-layer decoder-only transformer with a sparse MoE feed-forward backbone. Each token routes to 6 of 256 experts, plus 2 shared experts active on every token. Attention is a hybrid of local and global layers. The model is encoder-free and natively multimodal. Images are divided into 40×40-pixel patches and transformed using a four-layer hMLP. Audio is represented as dMel spectrograms. Both pass through a lightweight embedding layer and are processed jointly with text tokens. Numerics support covers BF16, MXFP8 and NVFP4. Audio input is WAV at 16 kHz, ideally under two minutes. Output is text only.
Inkling-Small began training after its larger counterpart. That let the research team revise the pre-training data mix and the machine learning recipe. The research team post-trained an earlier checkpoint, Inkling-Small (preview), in part using on-policy distillation with Inkling as the teacher. From that checkpoint, it continued scaling agentic coding RL for two weeks.
Benchmark results
The smaller model surpasses its teacher on reasoning and agentic coding. On Humanity’s Last Exam (text only) Inkling-Small scores 31.6%, ahead of Inkling’s 29.7%. SWE-bench Verified is 80.2% versus 77.6%, using a bash-only harness. Terminal-Bench 2.1 reaches 64.7% at best harness. Toolathlon Verified is 54.4%, against Inkling’s 45.5%. GPQA Diamond is 89.5%, AIME 2026 is 95.5%, and IFBench is 82.2%. ARC-AGI-2 rises to 40.1% from Inkling’s 36.5%.
SimpleQA Verified falls to 20.6% from Inkling’s 43.9%, and the AA Omniscience index drops to -9.0 from 2.1. Tau 3 Banking is 15.5% versus Inkling’s 23.7%. All evaluations ran at effort 0.99 and temperature 1.0, with a 256K max-token trajectory limit on coding evals. External scores are sourced from Artificial Analysis, Scale AI and ARC Prize.
Multimodality, epistemics and safety
Multimodal scores stay close to Inkling at lower cost. MMMU Pro is 74.0%. CharXiv RQ is 77.4%, rising to 81.3% when the model uses Python to crop, zoom and inspect charts programmatically. Audio MC is 54.9%, MMAU is 77.0%, and VoiceBench is 90.1%.
On epistemics, calibration was trained with RL against proper scoring rules on a large corpus of real-world forecasting questions. ForecastBench without search gives a Brier Index of 61.3 ± 0.46, ahead of Inkling’s 60.1 ± 0.54. On safety, StrongREJECT is 98.4%, FORTRESS adversarial is 71.6%, and FORTRESS benign is 96.9%. Thinking Machines Lab concluded the model presents no material uplift beyond the existing open-weight ecosystem. I
分享
阅读原文 ↗