The first response can be slower because Ollama loads the model into memory. The typing indicator means XLlama is waiting or receiving text; speed depends on your computer.
- Try a smaller downloaded model in Models.
- Test with a few cells and a short question. Large selections and long conversations require more work.
- Close other memory-intensive applications.
- In Settings → Keep model loaded, choose a longer period if you use the same model repeatedly.
- Use PROMPT or EXTRACT to reuse cached answers when the inputs match.
Use Stop to cancel a chat request or Cancel all AI requests for pending work. A cancelled chat that already produced answer text uses one Free response.
Run Help → Diagnostics to separate a connection problem from a slow model.
