FLUX
Image generation and editing models from Black Forest Labs, available as a hosted API or as downloadable weights you run on your own GPUs.
Open weights · 30 reviewed
Models whose weights you can pull and serve yourself, whatever the code situation. The licence is the thing to read here: several of the strongest releases carry community licences with revenue thresholds, attribution strings or field-of-use limits that a standard procurement review will not expect. Each entry names the licence and the catch.
Image generation and editing models from Black Forest Labs, available as a hosted API or as downloadable weights you run on your own GPUs.
OpenBMB's tokenizer-free TTS, Apache 2.0 on weights and code. VoxCPM2 is 2B parameters, 30 languages, 48kHz output and voice design from a written description with no reference clip.
Alibaba Qwen team's 20B MMDiT image foundation model plus its Edit and Layered variants, all Apache 2.0. Renders long-form Chinese and English text inside images at commercial quality.
Shanghai AI Lab's open vision-language family (书生·万象), nine sizes from 1B to 241B-A28B under Apache 2.0. The default open VLM for OCR, document parsing and GUI agent work.
MiniMax's open-weight flagship LLM (稀宇科技), distinct from the Hailuo video product: 428B mixture-of-experts, 23B active per token, 1M context and native image and video input.
Ant Group's open-weight MoE family (百灵, Bailing): Ling-3.0-flash at 124B total and 5.1B active, a 7.9B tiny variant and trillion-parameter Ring reasoning models, all MIT.
OpenBMB's on-device line (面壁小钢炮): MiniCPM5-1B at 1.08B parameters with a 131K context, plus MiniCPM-V 4.6 multimodal at 1.3B. Apache 2.0 on weights and code.
Open manipulation data and policy weights from AgiBot (智元机器人): a million-plus real robot trajectories, the GO-1 VLA and the Genie Envisioner world model, all non-commercial.
BAAI's embodied brain model (悟界·RoboBrain), Apache 2.0 at 4B and 8B, shipped in parallel NVIDIA and Moore Threads builds. It plans and points in 3D for a controller underneath.
Image and text to 3D from VAST in Beijing. Hosted Tripo 3.0 Ultra reaches 2 million polygons, while TripoSG 1.5B, TripoSR and UniRig ship as MIT-licensed weights you can run yourself.
RedNote's (小红书) MIT-licensed document parser. A single vision-language model does layout, reading order and text; dots.mocr at 3B added direct SVG output for charts in March 2026.
Alibaba's FunAudioLLM TTS line, Apache 2.0 on both code and weights. Fun-CosyVoice3-0.5B covers 9 languages and 18+ Chinese dialects with roughly 150ms bi-streaming latency.
Bilibili's zero-shot TTS with the finest emotion and duration control in the open field. IndexTTS-2.5 is 0.8B across five languages, but commercial use needs written bilibili authorization.
Alibaba Tongyi Lab's open-weight video generation family (通义万相 Wan). Apache 2.0 mixture-of-experts checkpoints you can download and fine-tune, separate from the closed Wan 3.0 API.
Alibaba's open-weights coding model family plus the Apache 2.0 Qwen Code terminal agent, separate artifacts from the hosted Qwen Chat product. Qwen3-Coder-Next runs 80B total, 3B active.
NVIDIA's open-weight family, 31.6B Nano up to the 550B Ultra MoE. Hybrid Mamba-Transformer, 1M context, weights plus roughly 3T tokens of training data published.
European lab publishing open weight models you can run on your own hardware. Large 3 is a 675B MoE with 41B active, 256K context and Apache 2.0 weights.
Tencent's open image-to-3D model with PBR texture synthesis and full training code. A separate artifact from the Hunyuan chat product, and not under an OSI-approved licence.
Shanghai AI Lab's open-weight family (书生·浦语), now the Intern-S scientific multimodal line: 35B up to 1T parameters, Apache 2.0, trained on domestic computing infrastructure.
Manycore Tech (群核科技) of Hangzhou, behind Kujiale and Coohom, open-sourced its indoor spatial stack. SpatialLM turns point clouds into structured layouts, SpatialGen generates scenes.
Google's open-weight family in five sizes, 2B effective up to 31B dense, text and image in with audio on the smaller ones, 256K context and Apache 2.0 weights since April 2026.
Meta's 30B dense multimodal agent model. Apache 2.0, 131K context, 4-bit checkpoints under 20GB and none of the Llama licence conditions attached.
Chinese AI lab producing open-source LLMs that rival proprietary models at a fraction of the cost, with strong reasoning and coding abilities.
Kai-Fu Lee’s open-source AI model family offering high-performance bilingual models with strong benchmark results across reasoning and coding tasks.
Baichuan AI’s open-source large language models focused on Chinese language understanding and enterprise AI applications.
Tencent’s large language model and AI platform, integrated across WeChat, Tencent Cloud, and the broader Tencent ecosystem.
hipu AI’s bilingual AI assistant powered by the GLM-4 model series, with strong academic roots from Tsinghua University and excellent bilingual capabilities
Moonshot AI’s long-context AI assistant, pioneering ultra-long document processing with a 2 million token context window.
Alibaba’s open-source large language model family and AI assistant, offering strong multilingual capabilities and deep enterprise integration with Alibaba Cloud.
Baidu’s flagship AI chatbot powered by ERNIE 4.5 and reasoning model ERNIE X1, offering multimodal capabilities across text, image, audio, and video.
Open weights is not the same as open source. A community licence can cap commercial use by revenue, require an attribution string in your interface or forbid whole categories of deployment. Where an entry carries those terms, they are stated in its cons.