Self-hosted
30 entities · 1 projects · 2 tools · 24 models · 3 news
Qwen3.8 27B
Qwen's dense 27B open-weight model for multilingual reasoning and generation.
Qwen3.8 Flash Next
A fast Qwen3.8 checkpoint built for efficient agent and assistant workloads.
Qwen3.6 35B-A3B
A Qwen mixture-of-experts checkpoint balancing total capacity with lower active parameters.
DeepSeek V4 Pro
DeepSeek's open V4 Pro checkpoint for advanced general reasoning and generation.
DeepSeek V4.1 Flash
An MIT-licensed DeepSeek V4.1 checkpoint for fast multimodal text generation.
DeepSeek V4 Flash
A faster open DeepSeek V4 checkpoint aimed at efficient inference.
Phi-4
Microsoft's compact open language model for reasoning and resource-conscious deployment.
Voxtral Mini Realtime
Mistral AI's compact realtime speech model for low-latency transcription workflows.
NVIDIA Nemotron 3 Super
NVIDIA's 120B mixture-of-experts Nemotron checkpoint for open reasoning workloads.
Llama 4 Scout
Meta's open-weight Llama 4 Scout mixture-of-experts instruction model.
Hugging Face Tokenizers
Hugging Face's Rust tokenizer library for research and production inference.
Phi-4 Mini Instruct
A smaller instruction-tuned Phi-4 checkpoint for local and edge-friendly workloads.
NVIDIA Nemotron 3.5 Lightning
An optimized 30B mixture-of-experts Nemotron release for accelerated inference.
Ministral 3 14B
A compact Mistral instruction model sized for controlled and local deployment.
Mistral Medium 3.5
Mistral AI's 128B-class open model for capable general-purpose inference.
Llama 4 Maverick
Meta's larger Llama 4 Maverick mixture-of-experts instruction checkpoint.
Mistral Small 4
A 119B mixture-of-experts Mistral checkpoint for efficient general workloads.
NVIDIA Nemotron 3 Embed 8B
An NVIDIA embedding model for retrieval, semantic search and enterprise indexing.
NVIDIA Kumo Tabular
NVIDIA's tabular foundation models predict classifications or numeric values from labeled examples, with three sizes spanning 28M to 215M parameters.
Fara 1.5 4B
A compact Microsoft research checkpoint focused on agentic computer-use tasks.
Holo4 27B
H Company's 27B agent model combines graphical interfaces, code, MCP and APIs; its downloadable checkpoint carries a noncommercial license.
ENZO
Self-hosted AI workspace with agents, skills and tools that runs on your own provider keys.
IBM Bob self-hosted deployment
IBM announced a self-hosted deployment option for its Bob AI software-development platform, including customer-managed and air-gapped environments.
Magnitude
Magnitude launched an open-source local inference engine for AI agents that tunes its kernels to a user's hardware.
oMLX
An Apache-2.0 local AI engine built around Apple's MLX ecosystem.
DeepSeek-R1
Reasoning model release that made open reasoning systems a central developer conversation.
Gemma 3
Google's open model family designed for capable, efficient multimodal applications.
GPT-OSS 120B
Open-weight reasoning model intended to run in environments controlled by developers.
Qwen3
Alibaba's open model family with reasoning and general-purpose variants for developers.
Open-weight models reset the deployment question
A short brief on why open-weight releases move model choice toward infrastructure design.