Ollama
ActifExécutez des modèles d'IA en local — le moteur d'inférence open source avec 176K étoiles GitHub
The default choice for local LLM inference — and for good reason. 176K GitHub stars, 9M+ users, MIT license, one-command setup. Runs on any hardware from Raspberry Pi 5 to dual H100s, with a model catalog spanning 4,500+ open-weight options. The OpenAI-compatible API with streaming, tool calling, structured outputs, and embeddings means no vendor lock-in — swap from Ollama to OpenAI (or vice versa) by changing one URL. Ollama Cloud (Pro $20/mo, Max $100/mo) extends the same surface to managed inference. The recent $65M raise confirms sustained investment. For any developer who wants to run AI locally — whether for privacy, cost control, or offline use — Ollama is where you start.
Each inference request is a stateless REST API call — no carried AI context between requests
Est-ce fait pour vous ?
Recommandé pour
- Développement IA local-first — exécutez plus de 4 500 modèles sur votre propre matériel sans coûts API
- Charges sensibles à la confidentialité — toute l'inférence reste sur votre infrastructure, ne quitte jamais la machine
- IA à coût contrôlé — le coût matériel fixe remplace la facturation par token ; point mort à environ $200/mois de dépenses API
- Environnements hors ligne/isolés — pas d'internet nécessaire après le pull du modèle
- Intégration chaîne d'outils développeur — API compatible OpenAI fonctionne avec LangChain, LlamaIndex, Hermes Agent, Continue.dev
Déconseillé pour
- Accès aux modèles frontière — ne peut exécuter que des modèles open-weight ; pas de GPT-5.6, Claude ou Gemini via Ollama
- Équipes sans matériel GPU — exécuter des modèles 70B+ sur CPU est impraticable (<1 tok/sec)
- Service managé zéro-opérations — les plans Cloud existent mais le produit local nécessite une gestion matérielle
Notre expérience
Tarifs
Open source- Run any compatible open-weight model on your own hardware
- No usage limits, no API keys, no data leaves your machine
- Daily quota for experimentation
- Same API surface as local runtime
- Full open-weight catalog
- Higher per-minute rate limits
- Run 3 cloud models at a time
- 50x more cloud usage than Free
- Upload and share private models
- Run 10 cloud models at a time
- 5x more usage than Pro
- New sign-ups paused for capacity
- High performance, up to 2x more than model gateways
- Zero data retention and logging
- Shared billing and administration
- Priority support
- SSO (coming soon)
Aucun changement de verdict pour l’instant
L’horloge tourne dès le premier jour : les changements apparaîtront ici à mesure que notre verdict évolue.
Sources
- Ollama — official websitejuil. 2026
- Ollama — GitHub repositoryjuil. 2026
- TechCrunch — Ollama raises $65M Series B, grows to nearly 9M usersjuil. 2026
- Pooya Golchian — Ollama Cloud Pricing 2026juil. 2026
- Thunder Compute — What is Ollama: Run AI Models Locally (July 2026)juil. 2026
- Kunal Ganglani — Best Local LLMs in 2026: Models, Hardware & Setup Guidejuil. 2026
- DanubeData — Run Ollama on a VPS: Self-Host Local LLMs in Europe (2026)juil. 2026
Journal de vérification
- Tarifs— Mis à jour
Agent automatisé
New Team plan: $25/seat/mo (5-seat minimum), includes shared billing, priority support. Pro $20/mo and Max $100/mo unchanged. Max sign-ups paused for capacity.
- Tarifs— Aucun changement
Agent automatisé
- Profil— Aucun changement
Importé au lancement
- Tarifs— Aucun changement
Importé au lancement