Skip to main content
Run large language models with Ollama, either locally or through Ollama Cloud. Ollama is a fantastic tool for running models both locally and in the cloud. Local Usage: Run models on your own hardware using the Ollama client. Cloud Usage: Access cloud-hosted models via Ollama Cloud with an API key. Ollama supports multiple open-source models. See the library here. Experiment with different models to find the best fit for your use case. Here are some general recommendations:
  • gpt-oss:120b-cloud is an excellent general-purpose cloud model for most tasks.
  • llama3.3 models are good for most basic use-cases.
  • qwen models perform specifically well with tool use.
  • deepseek-r1 models have strong reasoning capabilities.
  • phi4 models are powerful, while being really small in size.

Authentication (Ollama Cloud Only)

To use Ollama Cloud, set your OLLAMA_API_KEY environment variable. You can get an API key from Ollama Cloud.
When using Ollama Cloud, the host is automatically set to https://ollama.com. For local usage, no API key is required.

Set up a model

Local Usage

Install ollama and run a model:
run model
This starts an interactive session with the model. To download the model for use in an Agno agent:
pull model

Cloud Usage

For Ollama Cloud, no local Ollama server installation is required. Install the Ollama library, set up your API key as described in the Authentication section above, and access cloud-hosted models directly.

Examples

Local Usage

Once the model is available locally, use the Ollama model class to access it:

Cloud Usage

When using Ollama Cloud with an API key, the host is automatically set to https://ollama.com. You can omit the host parameter.
View more examples here.

Params

Ollama is a subclass of the Model class and has access to the same params.

Responses API

Ollama v0.13.3+ supports the OpenAI Responses API via the /v1/responses endpoint. Use OllamaResponses for this interface:
The Responses API is stateless. Each request is independent with no previous_response_id chaining. See OllamaResponses reference for full parameters.