Daily Tech Briefing
AI 科技速览
每天 5 分钟内学习 AI。获取最新的人工智能新闻,理解其重要性,并学习如何将其应用于您的工作。
Together AI · 2026/7/29 00:00:00

Together AI announces strategic partnership with Moonshot AI to natively serve Kimi models
AI 中文解读
国内大模型出海又有新动作!开源模型Kimi要和美国AI公司Together AI达成深度合作,以后Kimi系列模型将直接在对方平台上首发。这次合作的亮点是Kimi K3,一个参数规模高达2.8万亿的超级大模型,相当于开源界的“巨无霸”。它不仅能处理文本,还自带视觉能力,一次能“读”完100万token的内容,比很多同类产品都强悍。
通俗地说,以前人们想用顶级开源大模型,往往担心速度慢、部署难。这次合作让Kimi K3能在美国本土的算力上直接运行,中国企业用起来更流畅,数据也更安全。对于普通用户而言,最直观的感受是:以后用AI处理超长文件、编写复杂代码,或者让AI“看图说话”时,背后可选的免费模型技术含量一下子提高了。今后会有更多好用的AI应用能用上这种水平的模型,尤其国内开发者能借助它做出更聪明、更懂你的AI助手,还是值得期待的。
Today we're announcing a strategic partnership with Moonshot AI, one of the leading model labs pushing the open source frontier with large-scale MoE architectures. Under this partnership, Together AI becomes a launch platform for Moonshot's model releases, starting with Kimi K3 and extending to every open weights model Moonshot ships going forward.For developers building on open models, this means day zero access to Moonshot's frontier releases through Together AI’s US-hosted infrastructure, pricing model, and tooling, while ensuring zero data retention and compliant access to models and data. Developers can also post train these models with their own data to deliver the quality and performance for their app.Frontier performance, open weightsKimi K3 is the largest open model released to date: a 2.8T parameter sparse Mixture-of-Experts model with native vision support and a 1M token context window. Moonshot built it around two new architectural components:Kimi Delta Attention (KDA), which changes how information flows across sequence length and delivers significantly faster decoding at long context lengths.Attention Residuals (AttnRes), which improves how representations are retrieved across model depth, adding meaningful training efficiency at minimal extra compute cost.Layered on top are a set of optimizer refinements, including per-head Muon and quantile-based expert load balancing, that Moonshot says combine to deliver roughly 2.5x better scaling efficiency compared to Kimi K2. The result is a model built for long-horizon coding, agentic workflows, game development, and knowledge-intensive tasks, with benchmark results that put it in direct competition with the leading proprietary systems on the market today.Production Inference Platform Kimi models are available across Together AI Inference products – including Serverless, Provisioned Throughput and Dedicated Inference. Developers can go from trying these models to applying them to their production use cases with SLA-backed products and autoscaling.With Provisioned Throughput, teams can use reserved inference capacity for frontier open models with token-based pricing and a 99% uptime SLA. It’s the reserved-capacity guarantee developers already expect from closed-model providers, now available for open weights. No GPU-hour math, just guaranteed throughput at a predictable price. For teams looking for more control with all the benefits of a production-ready platform, they can use Dedicated Model Inference with fast deployment, better token economics and continuous research innovations shipped into the product.Post Training Developers can also post train Kimi models to deliver better quality for their target use case. With custom training, including full-weight and LoRA reinforcement learning as well as advanced supervised fine-tuning, developers can configure training through the Python SDK and granular primitives, and run multiple LoRA experiments concurrently on dedicated capacity. Custom training connects experimentation directly to production. When a checkpoint is ready to evaluate, it can be deployed natively to inference, with no separate handoff between training and serving and no rebuilding.What this partnership unlocks for developersDay zero availability: Kimi K3 is available on Together AI with the highest model quality and performance, validated by Moonshot. Future Moonshot releases will also ship on Together at launch.Proven scale with production-ready infrastructure: Together AI's research-optimized inference stack is tuned for large sparse MoE models like K3, so developers get the performance the architecture promises rather than a generic deployment. Together’s inference stack already serves production traffic for companies like Cursor, Y Combinator and Decagon, deploying coding and agentic workloads at scale. One integration, the full Moonshot lineup: As Moonshot ships new models, they'll land in the same Together AI Models library, behind the same API yo
分享
阅读原文 ↗