This is an automated email from the ASF dual-hosted git repository. davsclaus pushed a commit to branch quick-fix/jbang-docs-local-models in repository https://gitbox.apache.org/repos/asf/camel.git
commit 231e9d3ff629051d52e85474f979fbf8a5a5dd10 Author: Claus Ibsen <[email protected]> AuthorDate: Fri Oct 2 14:48:44 2026 +0200 chore: docs - camel-jbang local models and getting started get their own pages, and the terminal monitor overview gets shorter Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Signed-off-by: Claus Ibsen <[email protected]> --- docs/user-manual/modules/ROOT/nav.adoc | 2 + .../modules/ROOT/pages/camel-jbang-ai.adoc | 2 +- .../modules/ROOT/pages/camel-jbang-tui-ai.adoc | 208 +------------------- .../pages/camel-jbang-tui-getting-started.adoc | 124 ++++++++++++ .../ROOT/pages/camel-jbang-tui-local-models.adoc | 213 +++++++++++++++++++++ .../modules/ROOT/pages/camel-jbang-tui.adoc | 139 ++------------ 6 files changed, 361 insertions(+), 327 deletions(-) diff --git a/docs/user-manual/modules/ROOT/nav.adoc b/docs/user-manual/modules/ROOT/nav.adoc index f3a51c5d0dc2..08c67e8e6c96 100644 --- a/docs/user-manual/modules/ROOT/nav.adoc +++ b/docs/user-manual/modules/ROOT/nav.adoc @@ -11,10 +11,12 @@ *** xref:camel-jbang-kubernetes.adoc[Camel Kubernetes Plugin] *** xref:camel-jbang-test.adoc[Camel Testing Plugin] *** xref:camel-jbang-tui.adoc[Camel TUI] +**** xref:camel-jbang-tui-getting-started.adoc[Camel TUI Getting Started] **** xref:camel-jbang-tui-source-editor.adoc[Camel TUI Source Editor] **** xref:camel-jbang-tui-diagram.adoc[Camel TUI Diagram] **** xref:camel-jbang-tui-observe.adoc[Camel TUI Observing Integrations] **** xref:camel-jbang-tui-ai.adoc[Camel TUI AI Panel] +**** xref:camel-jbang-tui-local-models.adoc[Camel TUI Local Models] **** xref:camel-jbang-tui-ai-agents.adoc[Camel TUI and AI Agents] **** xref:camel-jbang-tui-settings.adoc[Camel TUI Actions, Settings and Themes] *** xref:camel-jbang-mcp.adoc[Camel MCP Server] diff --git a/docs/user-manual/modules/ROOT/pages/camel-jbang-ai.adoc b/docs/user-manual/modules/ROOT/pages/camel-jbang-ai.adoc index 2526cfb71d1c..8ce276952788 100644 --- a/docs/user-manual/modules/ROOT/pages/camel-jbang-ai.adoc +++ b/docs/user-manual/modules/ROOT/pages/camel-jbang-ai.adoc @@ -276,7 +276,7 @@ of reloading the model, and for a context window (`num_ctx`) chosen once per mod The tool-calling prompt alone is about 4.5k tokens, which is why 32k is the floor: Ollama's own default on machines with less than 24 GB is 4k and would truncate it. Keep all your Ollama clients on the same value; if you set `OLLAMA_CONTEXT_LENGTH` for the CLI, set it for `ollama serve` too. The TUI's Ollama -tab shows the window in use, and xref:camel-jbang-tui-ai.adoc#_working_with_a_local_ollama_model[Working +tab shows the window in use, and xref:camel-jbang-tui-local-models.adoc#_working_with_a_local_ollama_model[Working with a local Ollama model] explains how the AI panel manages its history inside it. === Using an OpenAI-compatible local server diff --git a/docs/user-manual/modules/ROOT/pages/camel-jbang-tui-ai.adoc b/docs/user-manual/modules/ROOT/pages/camel-jbang-tui-ai.adoc index 2a1ac2268d8b..33eb1244b06e 100644 --- a/docs/user-manual/modules/ROOT/pages/camel-jbang-tui-ai.adoc +++ b/docs/user-manual/modules/ROOT/pages/camel-jbang-tui-ai.adoc @@ -2,7 +2,7 @@ The AI panel (*F8*) lets you ask about your integrations in plain language: the AI sees what the TUI sees, reads the sources and the catalog, and can change files with your confirmation. It works with hosted models and with local -models through Ollama, so nothing has to leave your machine. +models through Ollama, so nothing has to leave your machine (see xref:camel-jbang-tui-local-models.adoc[Local Models]). See xref:camel-jbang-tui.adoc[Camel TUI] for getting started and the other pages. @@ -12,7 +12,7 @@ See xref:camel-jbang-tui.adoc[Camel TUI] for getting started and the other pages * <<_ai_panel_slash_commands,*Slash commands*>> -- run examples, change settings and steer the AI from the panel * <<_ai_project_overview,*Project overview*>> -- the AI explains the project once and the diagram shows it as capabilities * <<_ai_log,*AI log*>> -- what was asked, which tools ran and what they returned -* <<_ollama,*Ollama tab*>> -- the local models, their context window and what is loaded +* xref:camel-jbang-tui-local-models.adoc[*Local models*] -- Ollama or an OpenAI-compatible server on your own machine == Choosing an AI provider @@ -39,161 +39,9 @@ via *F2* -> _AI & MCP_ -> _Setup AI_. Use *F2* -> _Settings_ to pin a provider, image::jbang/camel-tui-ai-panel-answer.png[A local Ollama model explains why an order failed] -=== Using Ollama (local, no API key) - -Install Ollama natively for best performance — the native binary uses GPU acceleration -(Metal on macOS, CUDA/ROCm on Linux): - -[source,bash] ----- -# macOS -brew install ollama - -# Linux -curl -fsSL https://ollama.com/install.sh | sh - -# Pull a model — then open the TUI and press F8 -ollama pull qwen3.6:35b-a3b -camel tui ----- - -Ollama at `localhost:11434` is auto-detected. No configuration needed. - -IMPORTANT: The F8 AI panel works by invoking built-in tools to inspect your running Camel process. -Models smaller than ~14B do not reliably call tools and answer from training knowledge instead. -Use at least a 14B model. Prefer a mixture-of-experts model such as `qwen3.6:35b-a3b`: with only -3B parameters active per token it processes the tool-heavy prompt many times faster than a dense -27B/32B model, so answers start in seconds instead of a minute. - -*Models that work well* (tool-calling capable, ≥14B, default Q4_K_M quantization): - -[options="header"] -|=== -| Model | RAM | Notes -| `qwen3.6:35b-a3b` | ~23 GB | Recommended: fastest prompt processing, needs 32 GB+ -| `qwen2.5:14b` | ~9 GB | Minimum for 16 GB machines -| `qwen3.6:27b` | ~18 GB | Strong dense model, several times slower prompt processing -| `qwen2.5:32b` | ~20 GB | Good quality, slow prompt processing -| `hermes3:70b` | ~43 GB | Excellent tool calling, needs 64 GB+ -| `llama3.3:70b` | ~43 GB | Best open model, needs 64 GB+ -|=== - -NOTE: `camel infra run ollama` runs Ollama in Docker and bypasses GPU acceleration, -making inference significantly slower. Native install is preferred for development. - -NOTE: On Apple Silicon, use the default (GGUF) tags rather than the `-mlx` tags. The Ollama MLX -engine cannot yet reuse the cached prompt for Qwen 3.x models, so every question re-processes the -whole prompt, while the default engine reuses it and only processes what is new. - -=== Tool set for local models - -Every question sends the definitions of the `tui_*` tools the model may call, and a local model -pays for each of them in prompt-processing time. The panel therefore sends only the core set of -tools (state, tables, logs, errors, diagrams, topology, processor details, catalog docs, traces, -spans, route control, sending messages, source files, infra services, navigation, log level and -filters) to Ollama -and to any provider on `localhost`, which roughly halves the prompt. Hosted providers get every -tool, including the drawing, animation and automation tools. Use `/tools full` in the panel to send -all tools to a local model too, `/tools core` to trim the set for a hosted one, pick *AI Tools* in -*F2 -> Settings*, or set `camel.tui.ai.tools` in `.camel-cli.properties`. Ollama requests also ask the server to keep the -model loaded for 30 minutes and for a context window of 32k or 64k (see <<_working_with_a_local_ollama_model>>; -`OLLAMA_CONTEXT_LENGTH` overrides it), so follow-up questions reuse the cached prompt instead of reloading -the model. - -=== Working with a local Ollama model - -A local model is not a slower version of a hosted one; it spends its time differently, and the TUI -shows you where. This section explains what a question costs with Ollama and which knobs matter, -with figures measured on an Apple M4 Pro (64 GB) running `qwen3.6:35b-a3b`. - -*How a question is spent.* Ollama answers in three phases, and every timing the TUI shows maps to one -of them: - -1. *Load*: if the model is not in memory (first question, the keep-alive expired, or a request asked for - a different context size) Ollama starts a runner and loads the weights: 10 to 20 seconds for a 22 GB - model. The TUI asks Ollama to keep the model loaded for 30 minutes after each request. -2. *Prefill*: the prompt (system prompt, tool definitions, conversation history, your question) is - processed in one batch, at roughly 600 to 700 tokens per second when nothing is cached. -3. *Decode*: the answer is generated token by token, at 50 to 60 tokens per second for this model. - -The wait before anything appears, the time to first token, is load plus prefill. A cold first question -therefore takes 20 seconds before the first word; a warm one under a second. - -*One question, many requests.* The panel answers by calling the `tui_*` tools, and every tool call the -model makes costs a new request that sends the whole prompt again. A simple question can take 3 to 13 -requests. This is affordable only because Ollama caches the prompt prefix: the system prompt, the tool -definitions and the history are identical from one step to the next, so each step prefills only the new -tokens and takes about half a second. Across a session the cache hit is typically above 90%. The -Ollama tab shows the request count per question as `×N` and the cache hit per question; a question -with a high count and a short answer is the model exploring, which a smaller, sharper tool set reduces -(see <<_tool_set_for_local_models>>). - -*The context window.* The TUI decides the window it asks Ollama for (`num_ctx`) once per model: - -1. `OLLAMA_CONTEXT_LENGTH` in the environment wins when set. -2. If the model is already loaded, its window is adopted (raised to 32k when smaller), so the TUI never - makes Ollama reload a model another client is using. A request with a different `num_ctx` costs a - cold start and throws away everyone's prompt cache. -3. Otherwise 64k when the model's weights plus the KV cache of a 64k window fit the machine's memory, - computed from what `ollama show` reports (layers, KV heads, head size), else 32k. Overshooting is - worse than being conservative: Ollama then moves layers to the CPU and generation slows to a crawl. - -The static prefix of system prompt and core tools is about 4.5k tokens, and each question with tool -calls adds another 2k to 4k of history. The panel compacts the history once the prompt Ollama measured -for the last request passes half the window, capped at half of 64k even when a larger window was -adopted: every token of history is prefill time again after a cache loss (an idle unload after 30 -minutes, another client, a restart), and 64k at a few hundred tokens per second is already well over a -minute. Compacting rewrites the history, which invalidates Ollama's prompt cache, so the request after -a compaction prefills the whole prompt again: about 40 seconds for a 21k-token prompt. The panel -therefore compacts local history rarely and then thoroughly, down to about a quarter of the window, -and prints one line saying what it did and how long the next reply will take to start. `/compact` does -the same on demand with the same line; `/context` shows the window, the compaction point and the last -measured prompt; the panel title shows the fill as `ctx 34%`. - -*Choosing the model.* With no `camel.tui.ai.model` set the panel takes `llama3.2` when it is installed, -otherwise the first installed model; run `/model <name>` in the panel or set *AI Model* in -*F2 -> Settings* to pin one. The model must support tool calling and should have at least 14B -parameters; a mixture-of-experts model such as `qwen3.6:35b-a3b` prefills several times faster than a -dense model of similar quality, which is what matters for a tool-heavy prompt. *F2 -> Run Doctor* shows -whether Ollama was found, which models are installed and whether they are large enough. - -*Where to look.* The <<_ollama>> tab is the instrument for all of the above: tokens per second live and -per request, time to first token with cold starts marked, cache hit, how full the context window is -and how it grows per question, GPU and process load, and one line per question with the requests it -took, the question being answered right now on top with `working` as its reason and its time counting -up, and a footer with the average per question. In the AI panel, `/context` prints what the next request will cost, `/usage` and *Ctrl+U* the -session totals per question. - -*Remote and containerised Ollama.* Everything above applies to an Ollama on another host or inside -`camel infra run ollama` as well, with two differences: the container runs without GPU acceleration, -and the live runner state and host load on the Ollama tab need the server on the same machine. - -For the wider picture see the blog posts -link:/blog/2026/09/camel-local-model-benchmark/[We had a frontier AI coach a small local model through Camel] -on what a local model can do with Camel and what was changed to help it, and -link:/blog/2026/09/camel-genai-observability-jbang/[Observe Your Camel AI Routes with GenAI OpenTelemetry] -on observing routes that call Ollama. - -=== Using an OpenAI-compatible local server - -Set `LLM_API_KEY` and `LLM_BASE_URL` to connect to any OpenAI-compatible server -(LM Studio, vLLM, llama.cpp, GPT4All, …): - -[source,bash] ----- -export LLM_API_KEY=any-value -export LLM_BASE_URL=http://localhost:1234 -camel tui ----- - -`OPENAI_BASE_URL` is also accepted as an alternative to `LLM_BASE_URL`. - -The panel uses the first model the server lists on `/v1/models`. To use another one, run -`/model <name>` in the panel (`/model` alone lists what the server offers), set *AI Model* in -*F2 -> Settings*, or set `camel.tui.ai.model`. The model must support tool calling, otherwise the -panel answers from training data instead of inspecting your integration. When a request fails, the -panel shows the HTTP status and the server's error message, for example a model that the server -does not host. +To run the model on your own machine, with Ollama or an OpenAI-compatible server such as LM Studio, see +xref:camel-jbang-tui-local-models.adoc[Camel TUI Local Models]: which models work well, the tool set they get, +what a question costs and the Ollama tab. == AI panel slash commands @@ -228,7 +76,7 @@ cycles backward. | Show what the next request costs: provider and model, tool set, static prefix size, history size and the session total, and with Ollama the context window, the prompt size above which the history is compacted and the last measured prompt. Useful with local models, where prompt size is time. | `/compact` -| Shrink the conversation history sent to the model right away: older tool results are cut to their first lines and the oldest turns are dropped. With a hosted provider the panel does this automatically after each answer for all but the latest turn. With Ollama or another `localhost` provider it waits until the prompt Ollama measured passes half the context window (see <<_working_with_a_local_ollama_model>>), then compacts thoroughly, because a local server can reuse its cached prompt on [...] +| Shrink the conversation history sent to the model right away: older tool results are cut to their first lines and the oldest turns are dropped. With a hosted provider the panel does this automatically after each answer for all but the latest turn. With Ollama or another `localhost` provider it waits until the prompt Ollama measured passes half the context window (see xref:camel-jbang-tui-local-models.adoc#_working_with_a_local_ollama_model[Working with a local Ollama model]), then comp [...] | `/retry` | Send the last question again, starting from a clean turn in the model history. @@ -314,7 +162,7 @@ model called and the time spent in them, the number of requests, the tokens and, the context window got (for example `61.1s · 25 tool calls, limit reached (2.5s in tools) · 26 requests · 231.3k tokens · ctx 18%`). The panel allows 25 tool calls per question. When the model reaches that the byline says `limit reached`, the panel asks the model to answer from what it has found so far, and -the <<_ollama>> tab shows the question's request count in red with `limit` as its reason (the count is +the xref:camel-jbang-tui-local-models.adoc#_ollama_tab[Ollama tab] shows the question's request count in red with `limit` as its reason (the count is yellow from 10 requests, and the total time yellow from 30 seconds and orange from a minute). A question that runs into the limit is usually a tool that makes the model guess, not a weak model: the AI log shows which one. @@ -353,45 +201,3 @@ on a 10k-token prompt means the cache was hit, while a prefill of several second prompt was processed again. For OpenAI, Anthropic and Gemini it prints the number of `cached` input tokens the provider reported. Use it to check that follow-up questions are cheap before blaming the model for being slow. - -== Ollama - -The Ollama tab (under More, in the *AI* group) shows how the model served by a local -https://ollama.com[Ollama] is performing, in the spirit of an LLM dashboard. It works with or without a -running integration: it finds Ollama at `localhost:11434`, at the address of `camel infra run ollama`, -or at the endpoint the AI panel (*F8*) is using. The tab is listed only while an Ollama server answers; -the TUI checks every ten seconds, so it appears shortly after `ollama serve` starts. - -* *Model* -- the loaded model with its family, parameters, quantization, layers, experts (and how many - are active per token for a mixture-of-experts model), how much of it sits in GPU memory, the - allocated context length and when Ollama will unload it. With no model loaded, the installed models - are listed instead. -* *Throughput* -- decode and prefill tokens per second, *live* while the model is generating and - otherwise from the last request; time to first token and load time (a load of a second or more - is a cold start); session averages and a sparkline of the decode rate. -* *Context* -- how full the context window is, the share of the prompt served from Ollama's cache, - whether the model is working or idle, the speculative decoding method in use, and a per-turn trend - of how much of the window each AI panel prompt filled, with the session peak and the compactions seen. -* *Host* -- GPU utilization and memory (Apple silicon through `ioreg`, NVIDIA through `nvidia-smi`), - and CPU and memory of the Ollama server and its model runner. -* *Requests* -- one line per question asked in the AI panel (a question with tool calls costs one - request per step; *Enter* unfolds the steps) with the question text, prompt and generated tokens, - cache hit, the share of the context window reached, prefill and decode tokens per second, time to - first token, the time you waited and the stop reason. Calls made by Camel routes are listed as - their own lines. - -Two kinds of requests appear in the log. Questions asked in the AI panel with Ollama as the provider -come with the timings Ollama reports for each request (prompt evaluation, generation, model load, -total). Calls made by Camel routes through `camel-langchain4j-chat`, `camel-openai` or -`camel-spring-ai-chat` appear when the integration runs with GenAI observability (`--observe`, or -`--dep=camel:ai-observability`), tagged with the route id; Camel records tokens and duration for those, -not the phase split. - -The live figures, the context panel and the host panel need the model runner on the same machine: -Ollama starts a `llama-server` process per loaded model and the tab reads its slot state a few times a -second. Against a remote or containerised Ollama the tab keeps the model, per-request and session data -and says which panels are unavailable. - -Press *r* to reset the request log and the session totals, *F5* to refresh immediately. The same data -is available to AI agents through the `tui_get_ollama` MCP tool. For what the figures mean for the AI -panel and which knobs to turn, see <<_working_with_a_local_ollama_model>>. diff --git a/docs/user-manual/modules/ROOT/pages/camel-jbang-tui-getting-started.adoc b/docs/user-manual/modules/ROOT/pages/camel-jbang-tui-getting-started.adoc new file mode 100644 index 000000000000..f7f854f0f5e9 --- /dev/null +++ b/docs/user-manual/modules/ROOT/pages/camel-jbang-tui-getting-started.adoc @@ -0,0 +1,124 @@ += Camel TUI Getting Started + +There are several ways to get going with the TUI: point it at integrations that already run, run one of the +built-in examples, open a project directory, or connect an existing Spring Boot or Quarkus application. + +See xref:camel-jbang-tui.adoc[Camel TUI] for the tabs, the keyboard shortcuts and the other pages. + +== Option 1: Your Own Route + +Start a Camel integration in one terminal: + +[source,bash] +---- +camel run my-route.yaml +---- + +Open the TUI in another terminal: + +[source,bash] +---- +camel tui +---- + +The TUI auto-discovers every running Camel integration on your machine -- no configuration needed. + +== Option 2: Built-in Examples + +Don't have a route yet? The TUI ships with a catalog of ready-to-run examples. +Open the TUI and press *F2*, then select _Run an example_: + +[source,bash] +---- +camel tui +---- + +. Press *F2* to open the actions menu +. Select *Run an example...* +. Browse the example catalog -- type to filter by name +. Press *Enter* to launch the selected example + +Before launching, a run options form lets you choose the *runtime*: Camel Main (standalone), +Spring Boot, Quarkus, or JBang. This makes it easy to try any example on all three runtimes without +changing a single line of code. The first three run the example in a separate JVM that only +contains the dependencies of the example (like a production deployment), while JBang runs it +in-process in the Camel CLI JVM, which starts faster but has the CLI on the classpath as well. You can also set the integration name, toggle dev mode, and +add extra dependencies. When the port the app will listen on (the one you set, else 8080) is taken already, the form +says by whom, so a second app with an HTTP server does not fail to start. + +The example starts running in the background. The TUI auto-selects it as soon as it appears. +From there you can explore tabs, watch messages flow, inspect the route diagram, and experiment. + +If an example requires infrastructure (like Kafka or a database), the TUI automatically starts the +required Docker containers before launching the example. A notification in the footer shows the +progress. + +The same runtime selector is available in *Run from folder...* (F2 menu), which lets you point +the TUI at a local directory containing your routes. When a `pom.xml` is present, the runtime +is auto-detected and locked to match your project. + +TIP: Press *F1* or *?* on any screen for context-sensitive help. +Keyboard shortcuts are always shown in the footer bar. + +== Option 3: Open a Project Directory + +You can point the TUI directly at a project directory: + +[source,bash] +---- +camel tui . +camel tui /path/to/my-project +---- + +The TUI opens the directory in the Source tab so you can browse the project files immediately. +When a `pom.xml` is present, the runtime is auto-detected (Spring Boot, Quarkus, or Camel Main). +Press *F10* to run the project -- Maven projects are run with `camel run pom.xml`, which launches +them with the appropriate goal (`spring-boot:run`, `quarkus:dev`, or `camel:run`), and plain +directories are run with `camel run`. Either way the application logs to `~/.camel` so the Log tab +shows its logs. + +Until it runs, the project is listed as *Stopped*: the Overview says what it is (runtime, Maven or route files, its +folder) and the Source pane lists the routes it found and where they are, so *g* opens one. While it runs, the app +takes its place under the name it gives itself, with the project's name next to it; when the run ends, the project is +back as Stopped and still selected. + +This is a quick way to explore and run any Camel project without starting it separately first. + +image::jbang/camel-tui-main-quarkus-source.png[A Quarkus project in the Source tab: its Java route and the bean of the line] + +== Connecting Existing Spring Boot and Quarkus Applications + +The TUI auto-discovers integrations started with `camel run`. To monitor and control +existing Spring Boot or Quarkus applications, add the `camel-cli-connector` dependency +to your project. This is a lightweight runtime adapter that lets the TUI (and the Camel CLI) +communicate with your application -- all TUI features work the same way regardless of runtime. + +Spring Boot: + +[source,xml] +---- +<dependency> + <groupId>org.apache.camel.springboot</groupId> + <artifactId>camel-cli-connector-starter</artifactId> +</dependency> +---- + +Quarkus: + +[source,xml] +---- +<dependency> + <groupId>org.apache.camel.quarkus</groupId> + <artifactId>camel-quarkus-cli-connector</artifactId> +</dependency> +---- + +Once added, start your application normally and the TUI will discover it automatically. +No additional configuration is needed -- the connector auto-detects on the classpath and +registers the application with the local Camel CLI. + +TIP: When you open a Spring Boot, Quarkus or Camel Main project via *F2 > Open Project* and run +it with *F10*, the CLI connector dependency is automatically injected if it's not already in your +`pom.xml`. This means the TUI can monitor the application without modifying your project. + +See xref:camel-jbang-managing.adoc[Managing Integrations] for more details. diff --git a/docs/user-manual/modules/ROOT/pages/camel-jbang-tui-local-models.adoc b/docs/user-manual/modules/ROOT/pages/camel-jbang-tui-local-models.adoc new file mode 100644 index 000000000000..aacdf2def9bf --- /dev/null +++ b/docs/user-manual/modules/ROOT/pages/camel-jbang-tui-local-models.adoc @@ -0,0 +1,213 @@ += Camel TUI Local Models + +The xref:camel-jbang-tui-ai.adoc[AI panel] (*F8*) works with a model running on your own machine, through +https://ollama.com[Ollama] or any OpenAI-compatible server, so no API key is needed and nothing has to leave +your machine. A local model spends its time differently from a hosted one, and the TUI shows you where. + +See xref:camel-jbang-tui.adoc[Camel TUI] for getting started and the other pages. + +== Key Features + +* <<_using_ollama_local_no_api_key,*Ollama*>> -- auto-detected at `localhost:11434`, with the models that work well +* <<_using_an_openai_compatible_local_server,*OpenAI-compatible servers*>> -- LM Studio, vLLM, llama.cpp and others +* <<_tool_set_for_local_models,*A smaller tool set*>> -- local models get the core tools, which roughly halves the prompt +* <<_working_with_a_local_ollama_model,*What a question costs*>> -- load, prefill and decode, the prompt cache and the context window +* <<_ollama_tab,*Ollama tab*>> -- the loaded model, tokens per second, cache hit and context fill, live + +== Using Ollama (local, no API key) + +Install Ollama natively for best performance — the native binary uses GPU acceleration +(Metal on macOS, CUDA/ROCm on Linux): + +[source,bash] +---- +# macOS +brew install ollama + +# Linux +curl -fsSL https://ollama.com/install.sh | sh + +# Pull a model — then open the TUI and press F8 +ollama pull qwen3.6:35b-a3b +camel tui +---- + +Ollama at `localhost:11434` is auto-detected. No configuration needed. + +IMPORTANT: The F8 AI panel works by invoking built-in tools to inspect your running Camel process. +Models smaller than ~14B do not reliably call tools and answer from training knowledge instead. +Use at least a 14B model. Prefer a mixture-of-experts model such as `qwen3.6:35b-a3b`: with only +3B parameters active per token it processes the tool-heavy prompt many times faster than a dense +27B/32B model, so answers start in seconds instead of a minute. + +*Models that work well* (tool-calling capable, ≥14B, default Q4_K_M quantization): + +[options="header"] +|=== +| Model | RAM | Notes +| `qwen3.6:35b-a3b` | ~23 GB | Recommended: fastest prompt processing, needs 32 GB+ +| `qwen2.5:14b` | ~9 GB | Minimum for 16 GB machines +| `qwen3.6:27b` | ~18 GB | Strong dense model, several times slower prompt processing +| `qwen2.5:32b` | ~20 GB | Good quality, slow prompt processing +| `hermes3:70b` | ~43 GB | Excellent tool calling, needs 64 GB+ +| `llama3.3:70b` | ~43 GB | Best open model, needs 64 GB+ +|=== + +NOTE: `camel infra run ollama` runs Ollama in Docker and bypasses GPU acceleration, +making inference significantly slower. Native install is preferred for development. + +NOTE: On Apple Silicon, use the default (GGUF) tags rather than the `-mlx` tags. The Ollama MLX +engine cannot yet reuse the cached prompt for Qwen 3.x models, so every question re-processes the +whole prompt, while the default engine reuses it and only processes what is new. + +== Using an OpenAI-compatible local server + +Set `LLM_API_KEY` and `LLM_BASE_URL` to connect to any OpenAI-compatible server +(LM Studio, vLLM, llama.cpp, GPT4All, …): + +[source,bash] +---- +export LLM_API_KEY=any-value +export LLM_BASE_URL=http://localhost:1234 +camel tui +---- + +`OPENAI_BASE_URL` is also accepted as an alternative to `LLM_BASE_URL`. + +The panel uses the first model the server lists on `/v1/models`. To use another one, run +`/model <name>` in the panel (`/model` alone lists what the server offers), set *AI Model* in +*F2 -> Settings*, or set `camel.tui.ai.model`. The model must support tool calling, otherwise the +panel answers from training data instead of inspecting your integration. When a request fails, the +panel shows the HTTP status and the server's error message, for example a model that the server +does not host. + +== Tool set for local models + +Every question sends the definitions of the `tui_*` tools the model may call, and a local model +pays for each of them in prompt-processing time. The panel therefore sends only the core set of +tools (state, tables, logs, errors, diagrams, topology, processor details, catalog docs, traces, +spans, route control, sending messages, source files, infra services, navigation, log level and +filters) to Ollama +and to any provider on `localhost`, which roughly halves the prompt. Hosted providers get every +tool, including the drawing, animation and automation tools. Use `/tools full` in the panel to send +all tools to a local model too, `/tools core` to trim the set for a hosted one, pick *AI Tools* in +*F2 -> Settings*, or set `camel.tui.ai.tools` in `.camel-cli.properties`. Ollama requests also ask the server to keep the +model loaded for 30 minutes and for a context window of 32k or 64k (see <<_working_with_a_local_ollama_model>>; +`OLLAMA_CONTEXT_LENGTH` overrides it), so follow-up questions reuse the cached prompt instead of reloading +the model. + +== Working with a local Ollama model + +A local model is not a slower version of a hosted one; it spends its time differently, and the TUI +shows you where. This section explains what a question costs with Ollama and which knobs matter, +with figures measured on an Apple M4 Pro (64 GB) running `qwen3.6:35b-a3b`. + +*How a question is spent.* Ollama answers in three phases, and every timing the TUI shows maps to one +of them: + +1. *Load*: if the model is not in memory (first question, the keep-alive expired, or a request asked for + a different context size) Ollama starts a runner and loads the weights: 10 to 20 seconds for a 22 GB + model. The TUI asks Ollama to keep the model loaded for 30 minutes after each request. +2. *Prefill*: the prompt (system prompt, tool definitions, conversation history, your question) is + processed in one batch, at roughly 600 to 700 tokens per second when nothing is cached. +3. *Decode*: the answer is generated token by token, at 50 to 60 tokens per second for this model. + +The wait before anything appears, the time to first token, is load plus prefill. A cold first question +therefore takes 20 seconds before the first word; a warm one under a second. + +*One question, many requests.* The panel answers by calling the `tui_*` tools, and every tool call the +model makes costs a new request that sends the whole prompt again. A simple question can take 3 to 13 +requests. This is affordable only because Ollama caches the prompt prefix: the system prompt, the tool +definitions and the history are identical from one step to the next, so each step prefills only the new +tokens and takes about half a second. Across a session the cache hit is typically above 90%. The +Ollama tab shows the request count per question as `×N` and the cache hit per question; a question +with a high count and a short answer is the model exploring, which a smaller, sharper tool set reduces +(see <<_tool_set_for_local_models>>). + +*The context window.* The TUI decides the window it asks Ollama for (`num_ctx`) once per model: + +1. `OLLAMA_CONTEXT_LENGTH` in the environment wins when set. +2. If the model is already loaded, its window is adopted (raised to 32k when smaller), so the TUI never + makes Ollama reload a model another client is using. A request with a different `num_ctx` costs a + cold start and throws away everyone's prompt cache. +3. Otherwise 64k when the model's weights plus the KV cache of a 64k window fit the machine's memory, + computed from what `ollama show` reports (layers, KV heads, head size), else 32k. Overshooting is + worse than being conservative: Ollama then moves layers to the CPU and generation slows to a crawl. + +The static prefix of system prompt and core tools is about 4.5k tokens, and each question with tool +calls adds another 2k to 4k of history. The panel compacts the history once the prompt Ollama measured +for the last request passes half the window, capped at half of 64k even when a larger window was +adopted: every token of history is prefill time again after a cache loss (an idle unload after 30 +minutes, another client, a restart), and 64k at a few hundred tokens per second is already well over a +minute. Compacting rewrites the history, which invalidates Ollama's prompt cache, so the request after +a compaction prefills the whole prompt again: about 40 seconds for a 21k-token prompt. The panel +therefore compacts local history rarely and then thoroughly, down to about a quarter of the window, +and prints one line saying what it did and how long the next reply will take to start. `/compact` does +the same on demand with the same line; `/context` shows the window, the compaction point and the last +measured prompt; the panel title shows the fill as `ctx 34%`. + +*Choosing the model.* With no `camel.tui.ai.model` set the panel takes `llama3.2` when it is installed, +otherwise the first installed model; run `/model <name>` in the panel or set *AI Model* in +*F2 -> Settings* to pin one. The model must support tool calling and should have at least 14B +parameters; a mixture-of-experts model such as `qwen3.6:35b-a3b` prefills several times faster than a +dense model of similar quality, which is what matters for a tool-heavy prompt. *F2 -> Run Doctor* shows +whether Ollama was found, which models are installed and whether they are large enough. + +*Where to look.* The <<_ollama_tab,Ollama tab>> is the instrument for all of the above: tokens per second live and +per request, time to first token with cold starts marked, cache hit, how full the context window is +and how it grows per question, GPU and process load, and one line per question with the requests it +took, the question being answered right now on top with `working` as its reason and its time counting +up, and a footer with the average per question. In the AI panel, `/context` prints what the next request will cost, `/usage` and *Ctrl+U* the +session totals per question. + +*Remote and containerised Ollama.* Everything above applies to an Ollama on another host or inside +`camel infra run ollama` as well, with two differences: the container runs without GPU acceleration, +and the live runner state and host load on the Ollama tab need the server on the same machine. + +For the wider picture see the blog posts +link:/blog/2026/09/camel-local-model-benchmark/[We had a frontier AI coach a small local model through Camel] +on what a local model can do with Camel and what was changed to help it, and +link:/blog/2026/09/camel-genai-observability-jbang/[Observe Your Camel AI Routes with GenAI OpenTelemetry] +on observing routes that call Ollama. + +== Ollama tab + +The Ollama tab (under More, in the *AI* group) shows how the model served by a local +https://ollama.com[Ollama] is performing, in the spirit of an LLM dashboard. It works with or without a +running integration: it finds Ollama at `localhost:11434`, at the address of `camel infra run ollama`, +or at the endpoint the AI panel (*F8*) is using. The tab is listed only while an Ollama server answers; +the TUI checks every ten seconds, so it appears shortly after `ollama serve` starts. + +* *Model* -- the loaded model with its family, parameters, quantization, layers, experts (and how many + are active per token for a mixture-of-experts model), how much of it sits in GPU memory, the + allocated context length and when Ollama will unload it. With no model loaded, the installed models + are listed instead. +* *Throughput* -- decode and prefill tokens per second, *live* while the model is generating and + otherwise from the last request; time to first token and load time (a load of a second or more + is a cold start); session averages and a sparkline of the decode rate. +* *Context* -- how full the context window is, the share of the prompt served from Ollama's cache, + whether the model is working or idle, the speculative decoding method in use, and a per-turn trend + of how much of the window each AI panel prompt filled, with the session peak and the compactions seen. +* *Host* -- GPU utilization and memory (Apple silicon through `ioreg`, NVIDIA through `nvidia-smi`), + and CPU and memory of the Ollama server and its model runner. +* *Requests* -- one line per question asked in the AI panel (a question with tool calls costs one + request per step; *Enter* unfolds the steps) with the question text, prompt and generated tokens, + cache hit, the share of the context window reached, prefill and decode tokens per second, time to + first token, the time you waited and the stop reason. Calls made by Camel routes are listed as + their own lines. + +Two kinds of requests appear in the log. Questions asked in the AI panel with Ollama as the provider +come with the timings Ollama reports for each request (prompt evaluation, generation, model load, +total). Calls made by Camel routes through `camel-langchain4j-chat`, `camel-openai` or +`camel-spring-ai-chat` appear when the integration runs with GenAI observability (`--observe`, or +`--dep=camel:ai-observability`), tagged with the route id; Camel records tokens and duration for those, +not the phase split. + +The live figures, the context panel and the host panel need the model runner on the same machine: +Ollama starts a `llama-server` process per loaded model and the tab reads its slot state a few times a +second. Against a remote or containerised Ollama the tab keeps the model, per-request and session data +and says which panels are unavailable. + +Press *r* to reset the request log and the session totals, *F5* to refresh immediately. The same data +is available to AI agents through the `tui_get_ollama` MCP tool. For what the figures mean for the AI +panel and which knobs to turn, see <<_working_with_a_local_ollama_model>>. diff --git a/docs/user-manual/modules/ROOT/pages/camel-jbang-tui.adoc b/docs/user-manual/modules/ROOT/pages/camel-jbang-tui.adoc index 2bb536b6a25d..78f90dddee3e 100644 --- a/docs/user-manual/modules/ROOT/pages/camel-jbang-tui.adoc +++ b/docs/user-manual/modules/ROOT/pages/camel-jbang-tui.adoc @@ -14,7 +14,7 @@ image::jbang/camel-tui-overview.png[TUI Overview showing multiple routes] == Key Features * *Run anything* -- your own routes, the built-in examples, or an existing Spring Boot or Quarkus project - (<<_getting_started,Getting started>>). + (xref:camel-jbang-tui-getting-started.adoc[Getting started]). * *Write routes with help* -- a xref:camel-jbang-tui-source-editor.adoc[source editor] that checks YAML, Java and XML routes as you type, completes endpoint options, fixes problems and shows the live run data of each line. * *See the integration* -- the xref:camel-jbang-tui-diagram.adoc[diagram] from the architecture down to one route, @@ -28,126 +28,23 @@ image::jbang/camel-tui-overview.png[TUI Overview showing multiple routes] == Getting Started -You can start using the TUI in two ways: with your own route, or by running one of the built-in examples. - -=== Option 1: Your Own Route - -Start a Camel integration in one terminal: - [source,bash] ---- -camel run my-route.yaml ----- - -Open the TUI in another terminal: - -[source,bash] ----- -camel tui ----- - -The TUI auto-discovers every running Camel integration on your machine -- no configuration needed. - -=== Option 2: Built-in Examples +camel run my-route.yaml # in one terminal +camel tui # in another: auto-discovers every running integration -Don't have a route yet? The TUI ships with a catalog of ready-to-run examples. -Open the TUI and press *F2*, then select _Run an example_: - -[source,bash] ----- -camel tui +camel tui . # or open a project directory and press F10 to run it ---- -. Press *F2* to open the actions menu -. Select *Run an example...* -. Browse the example catalog -- type to filter by name -. Press *Enter* to launch the selected example - -Before launching, a run options form lets you choose the *runtime*: Camel Main (standalone), -Spring Boot, Quarkus, or JBang. This makes it easy to try any example on all three runtimes without -changing a single line of code. The first three run the example in a separate JVM that only -contains the dependencies of the example (like a production deployment), while JBang runs it -in-process in the Camel CLI JVM, which starts faster but has the CLI on the classpath as well. You can also set the integration name, toggle dev mode, and -add extra dependencies. When the port the app will listen on (the one you set, else 8080) is taken already, the form -says by whom, so a second app with an HTTP server does not fail to start. - -The example starts running in the background. The TUI auto-selects it as soon as it appears. -From there you can explore tabs, watch messages flow, inspect the route diagram, and experiment. - -If an example requires infrastructure (like Kafka or a database), the TUI automatically starts the -required Docker containers before launching the example. A notification in the footer shows the -progress. +No route yet? Press *F2* and pick _Run an example..._ to run one of the built-in examples on Camel Main, +Spring Boot, Quarkus or JBang. -The same runtime selector is available in *Run from folder...* (F2 menu), which lets you point -the TUI at a local directory containing your routes. When a `pom.xml` is present, the runtime -is auto-detected and locked to match your project. +See xref:camel-jbang-tui-getting-started.adoc[Camel TUI Getting Started] for each way to start, and for +connecting existing Spring Boot and Quarkus applications. TIP: Press *F1* or *?* on any screen for context-sensitive help. Keyboard shortcuts are always shown in the footer bar. -=== Option 3: Open a Project Directory - -You can point the TUI directly at a project directory: - -[source,bash] ----- -camel tui . -camel tui /path/to/my-project ----- - -The TUI opens the directory in the Source tab so you can browse the project files immediately. -When a `pom.xml` is present, the runtime is auto-detected (Spring Boot, Quarkus, or Camel Main). -Press *F10* to run the project -- Maven projects are run with `camel run pom.xml`, which launches -them with the appropriate goal (`spring-boot:run`, `quarkus:dev`, or `camel:run`), and plain -directories are run with `camel run`. Either way the application logs to `~/.camel` so the Log tab -shows its logs. - -Until it runs, the project is listed as *Stopped*: the Overview says what it is (runtime, Maven or route files, its -folder) and the Source pane lists the routes it found and where they are, so *g* opens one. While it runs, the app -takes its place under the name it gives itself, with the project's name next to it; when the run ends, the project is -back as Stopped and still selected. - -This is a quick way to explore and run any Camel project without starting it separately first. - -image::jbang/camel-tui-main-quarkus-source.png[A Quarkus project in the Source tab: its Java route and the bean of the line] - -=== Connecting Existing Spring Boot and Quarkus Applications - -The TUI auto-discovers integrations started with `camel run`. To monitor and control -existing Spring Boot or Quarkus applications, add the `camel-cli-connector` dependency -to your project. This is a lightweight runtime adapter that lets the TUI (and the Camel CLI) -communicate with your application -- all TUI features work the same way regardless of runtime. - -Spring Boot: - -[source,xml] ----- -<dependency> - <groupId>org.apache.camel.springboot</groupId> - <artifactId>camel-cli-connector-starter</artifactId> -</dependency> ----- - -Quarkus: - -[source,xml] ----- -<dependency> - <groupId>org.apache.camel.quarkus</groupId> - <artifactId>camel-quarkus-cli-connector</artifactId> -</dependency> ----- - -Once added, start your application normally and the TUI will discover it automatically. -No additional configuration is needed -- the connector auto-detects on the classpath and -registers the application with the local Camel CLI. - -TIP: When you open a Spring Boot, Quarkus or Camel Main project via *F2 > Open Project* and run -it with *F10*, the CLI connector dependency is automatically injected if it's not already in your -`pom.xml`. This means the TUI can monitor the application without modifying your project. - -See xref:camel-jbang-managing.adoc[Managing Integrations] for more details. - == Tabs Overview The TUI organizes information into tabs. Press number keys *1* through *0* to jump directly @@ -190,30 +87,22 @@ Two panels can be opened on top of any tab: *F6* opens an xref:camel-jbang-tui-s for running `camel` commands, and *F8* opens the xref:camel-jbang-tui-ai.adoc[AI panel] for asking questions about the running integrations. -== Source Code Browser - -The Source tab (Tab 2) is a file explorer and editor for your project code that knows Camel. It checks your YAML, -Java and XML routes as you type, completes endpoint options from the Camel catalog, explains the line the cursor is on, -fixes problems for you (or asks the AI to), and shows the exchanges and failures of each line of the running -integration next to the code. - -image::jbang/camel-tui-source-live-run-data.png[Live run data in the source editor] - -See xref:camel-jbang-tui-source-editor.adoc[Camel TUI Source Editor] for all its features and keys. - == More about the TUI [cols="1,3",options="header"] |=== | Page | What it covers +| xref:camel-jbang-tui-getting-started.adoc[Getting Started] | Your own routes, the built-in examples, opening a project, +connecting Spring Boot and Quarkus applications | xref:camel-jbang-tui-source-editor.adoc[Source Editor] | Reading and writing routes: checks as you type, quick fixes, fix with AI, completion, quick documentation, navigation, live run data | xref:camel-jbang-tui-diagram.adoc[Diagram] | The architecture, topology and route views, external endpoints, metrics | xref:camel-jbang-tui-observe.adoc[Observing Integrations] | Activity, message history, errors, spans, process, HTTP probe, CVE audit, Kafka, SQL, memory leaks, JFR, catalog -| xref:camel-jbang-tui-ai.adoc[AI Panel] | AI providers and local models, slash commands, project overview, AI log, the -Ollama tab +| xref:camel-jbang-tui-ai.adoc[AI Panel] | AI providers, slash commands, project overview, AI log +| xref:camel-jbang-tui-local-models.adoc[Local Models] | Ollama and OpenAI-compatible servers, the tool set for local +models, what a question costs, the Ollama tab | xref:camel-jbang-tui-ai-agents.adoc[AI Agents] | MCP and ACP: what an AI agent can see and do, edits you confirm | xref:camel-jbang-tui-settings.adoc[Actions, Settings and Themes] | The actions menu, embedded shell, themes, settings, browser access, recording demos @@ -276,7 +165,7 @@ See the xref:camel-jbang-tui-source-editor.adoc#_keyboard_shortcuts[keyboard sho | `100` | `--theme` -| Color theme for this session (e.g., `dark`, `tokyo-night`, `dracula`). See <<Theme>> for the full list of 21 themes. Overrides the persisted `camel.tui.theme` preference when set. +| Color theme for this session (e.g., `dark`, `tokyo-night`, `dracula`). See xref:camel-jbang-tui-settings.adoc#_theme[Theme] for the full list of 21 themes. Overrides the persisted `camel.tui.theme` preference when set. | | `--record`
