This is an automated email from the ASF dual-hosted git repository.
davsclaus pushed a commit to branch main
in repository https://gitbox.apache.org/repos/asf/camel.git
The following commit(s) were added to refs/heads/main by this push:
new 4062f82648b4 chore: docs - camel-jbang local models and getting
started get their own pages (#27282)
4062f82648b4 is described below
commit 4062f82648b484736219c804141f668a6ec66bcb
Author: Claus Ibsen <[email protected]>
AuthorDate: Fri Oct 2 15:06:45 2026 +0200
chore: docs - camel-jbang local models and getting started get their own
pages (#27282)
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
---
docs/user-manual/modules/ROOT/nav.adoc | 2 +
.../modules/ROOT/pages/camel-jbang-ai.adoc | 2 +-
.../modules/ROOT/pages/camel-jbang-tui-ai.adoc | 208 +-------------------
.../pages/camel-jbang-tui-getting-started.adoc | 124 ++++++++++++
.../ROOT/pages/camel-jbang-tui-local-models.adoc | 213 +++++++++++++++++++++
.../modules/ROOT/pages/camel-jbang-tui.adoc | 139 ++------------
6 files changed, 361 insertions(+), 327 deletions(-)
diff --git a/docs/user-manual/modules/ROOT/nav.adoc
b/docs/user-manual/modules/ROOT/nav.adoc
index f3a51c5d0dc2..08c67e8e6c96 100644
--- a/docs/user-manual/modules/ROOT/nav.adoc
+++ b/docs/user-manual/modules/ROOT/nav.adoc
@@ -11,10 +11,12 @@
*** xref:camel-jbang-kubernetes.adoc[Camel Kubernetes Plugin]
*** xref:camel-jbang-test.adoc[Camel Testing Plugin]
*** xref:camel-jbang-tui.adoc[Camel TUI]
+**** xref:camel-jbang-tui-getting-started.adoc[Camel TUI Getting Started]
**** xref:camel-jbang-tui-source-editor.adoc[Camel TUI Source Editor]
**** xref:camel-jbang-tui-diagram.adoc[Camel TUI Diagram]
**** xref:camel-jbang-tui-observe.adoc[Camel TUI Observing Integrations]
**** xref:camel-jbang-tui-ai.adoc[Camel TUI AI Panel]
+**** xref:camel-jbang-tui-local-models.adoc[Camel TUI Local Models]
**** xref:camel-jbang-tui-ai-agents.adoc[Camel TUI and AI Agents]
**** xref:camel-jbang-tui-settings.adoc[Camel TUI Actions, Settings and Themes]
*** xref:camel-jbang-mcp.adoc[Camel MCP Server]
diff --git a/docs/user-manual/modules/ROOT/pages/camel-jbang-ai.adoc
b/docs/user-manual/modules/ROOT/pages/camel-jbang-ai.adoc
index 2526cfb71d1c..8ce276952788 100644
--- a/docs/user-manual/modules/ROOT/pages/camel-jbang-ai.adoc
+++ b/docs/user-manual/modules/ROOT/pages/camel-jbang-ai.adoc
@@ -276,7 +276,7 @@ of reloading the model, and for a context window
(`num_ctx`) chosen once per mod
The tool-calling prompt alone is about 4.5k tokens, which is why 32k is the
floor: Ollama's own default
on machines with less than 24 GB is 4k and would truncate it. Keep all your
Ollama clients on the same
value; if you set `OLLAMA_CONTEXT_LENGTH` for the CLI, set it for `ollama
serve` too. The TUI's Ollama
-tab shows the window in use, and
xref:camel-jbang-tui-ai.adoc#_working_with_a_local_ollama_model[Working
+tab shows the window in use, and
xref:camel-jbang-tui-local-models.adoc#_working_with_a_local_ollama_model[Working
with a local Ollama model] explains how the AI panel manages its history
inside it.
=== Using an OpenAI-compatible local server
diff --git a/docs/user-manual/modules/ROOT/pages/camel-jbang-tui-ai.adoc
b/docs/user-manual/modules/ROOT/pages/camel-jbang-tui-ai.adoc
index 2a1ac2268d8b..33eb1244b06e 100644
--- a/docs/user-manual/modules/ROOT/pages/camel-jbang-tui-ai.adoc
+++ b/docs/user-manual/modules/ROOT/pages/camel-jbang-tui-ai.adoc
@@ -2,7 +2,7 @@
The AI panel (*F8*) lets you ask about your integrations in plain language:
the AI sees what the TUI sees, reads the
sources and the catalog, and can change files with your confirmation. It works
with hosted models and with local
-models through Ollama, so nothing has to leave your machine.
+models through Ollama, so nothing has to leave your machine (see
xref:camel-jbang-tui-local-models.adoc[Local Models]).
See xref:camel-jbang-tui.adoc[Camel TUI] for getting started and the other
pages.
@@ -12,7 +12,7 @@ See xref:camel-jbang-tui.adoc[Camel TUI] for getting started
and the other pages
* <<_ai_panel_slash_commands,*Slash commands*>> -- run examples, change
settings and steer the AI from the panel
* <<_ai_project_overview,*Project overview*>> -- the AI explains the project
once and the diagram shows it as capabilities
* <<_ai_log,*AI log*>> -- what was asked, which tools ran and what they
returned
-* <<_ollama,*Ollama tab*>> -- the local models, their context window and what
is loaded
+* xref:camel-jbang-tui-local-models.adoc[*Local models*] -- Ollama or an
OpenAI-compatible server on your own machine
== Choosing an AI provider
@@ -39,161 +39,9 @@ via *F2* -> _AI & MCP_ -> _Setup AI_. Use *F2* ->
_Settings_ to pin a provider,
image::jbang/camel-tui-ai-panel-answer.png[A local Ollama model explains why
an order failed]
-=== Using Ollama (local, no API key)
-
-Install Ollama natively for best performance — the native binary uses GPU
acceleration
-(Metal on macOS, CUDA/ROCm on Linux):
-
-[source,bash]
-----
-# macOS
-brew install ollama
-
-# Linux
-curl -fsSL https://ollama.com/install.sh | sh
-
-# Pull a model — then open the TUI and press F8
-ollama pull qwen3.6:35b-a3b
-camel tui
-----
-
-Ollama at `localhost:11434` is auto-detected. No configuration needed.
-
-IMPORTANT: The F8 AI panel works by invoking built-in tools to inspect your
running Camel process.
-Models smaller than ~14B do not reliably call tools and answer from training
knowledge instead.
-Use at least a 14B model. Prefer a mixture-of-experts model such as
`qwen3.6:35b-a3b`: with only
-3B parameters active per token it processes the tool-heavy prompt many times
faster than a dense
-27B/32B model, so answers start in seconds instead of a minute.
-
-*Models that work well* (tool-calling capable, ≥14B, default Q4_K_M
quantization):
-
-[options="header"]
-|===
-| Model | RAM | Notes
-| `qwen3.6:35b-a3b` | ~23 GB | Recommended: fastest prompt processing, needs
32 GB+
-| `qwen2.5:14b` | ~9 GB | Minimum for 16 GB machines
-| `qwen3.6:27b` | ~18 GB | Strong dense model, several times slower prompt
processing
-| `qwen2.5:32b` | ~20 GB | Good quality, slow prompt processing
-| `hermes3:70b` | ~43 GB | Excellent tool calling, needs 64 GB+
-| `llama3.3:70b` | ~43 GB | Best open model, needs 64 GB+
-|===
-
-NOTE: `camel infra run ollama` runs Ollama in Docker and bypasses GPU
acceleration,
-making inference significantly slower. Native install is preferred for
development.
-
-NOTE: On Apple Silicon, use the default (GGUF) tags rather than the `-mlx`
tags. The Ollama MLX
-engine cannot yet reuse the cached prompt for Qwen 3.x models, so every
question re-processes the
-whole prompt, while the default engine reuses it and only processes what is
new.
-
-=== Tool set for local models
-
-Every question sends the definitions of the `tui_*` tools the model may call,
and a local model
-pays for each of them in prompt-processing time. The panel therefore sends
only the core set of
-tools (state, tables, logs, errors, diagrams, topology, processor details,
catalog docs, traces,
-spans, route control, sending messages, source files, infra services,
navigation, log level and
-filters) to Ollama
-and to any provider on `localhost`, which roughly halves the prompt. Hosted
providers get every
-tool, including the drawing, animation and automation tools. Use `/tools full`
in the panel to send
-all tools to a local model too, `/tools core` to trim the set for a hosted
one, pick *AI Tools* in
-*F2 -> Settings*, or set `camel.tui.ai.tools` in `.camel-cli.properties`.
Ollama requests also ask the server to keep the
-model loaded for 30 minutes and for a context window of 32k or 64k (see
<<_working_with_a_local_ollama_model>>;
-`OLLAMA_CONTEXT_LENGTH` overrides it), so follow-up questions reuse the cached
prompt instead of reloading
-the model.
-
-=== Working with a local Ollama model
-
-A local model is not a slower version of a hosted one; it spends its time
differently, and the TUI
-shows you where. This section explains what a question costs with Ollama and
which knobs matter,
-with figures measured on an Apple M4 Pro (64 GB) running `qwen3.6:35b-a3b`.
-
-*How a question is spent.* Ollama answers in three phases, and every timing
the TUI shows maps to one
-of them:
-
-1. *Load*: if the model is not in memory (first question, the keep-alive
expired, or a request asked for
- a different context size) Ollama starts a runner and loads the weights: 10
to 20 seconds for a 22 GB
- model. The TUI asks Ollama to keep the model loaded for 30 minutes after
each request.
-2. *Prefill*: the prompt (system prompt, tool definitions, conversation
history, your question) is
- processed in one batch, at roughly 600 to 700 tokens per second when
nothing is cached.
-3. *Decode*: the answer is generated token by token, at 50 to 60 tokens per
second for this model.
-
-The wait before anything appears, the time to first token, is load plus
prefill. A cold first question
-therefore takes 20 seconds before the first word; a warm one under a second.
-
-*One question, many requests.* The panel answers by calling the `tui_*` tools,
and every tool call the
-model makes costs a new request that sends the whole prompt again. A simple
question can take 3 to 13
-requests. This is affordable only because Ollama caches the prompt prefix: the
system prompt, the tool
-definitions and the history are identical from one step to the next, so each
step prefills only the new
-tokens and takes about half a second. Across a session the cache hit is
typically above 90%. The
-Ollama tab shows the request count per question as `×N` and the cache hit per
question; a question
-with a high count and a short answer is the model exploring, which a smaller,
sharper tool set reduces
-(see <<_tool_set_for_local_models>>).
-
-*The context window.* The TUI decides the window it asks Ollama for
(`num_ctx`) once per model:
-
-1. `OLLAMA_CONTEXT_LENGTH` in the environment wins when set.
-2. If the model is already loaded, its window is adopted (raised to 32k when
smaller), so the TUI never
- makes Ollama reload a model another client is using. A request with a
different `num_ctx` costs a
- cold start and throws away everyone's prompt cache.
-3. Otherwise 64k when the model's weights plus the KV cache of a 64k window
fit the machine's memory,
- computed from what `ollama show` reports (layers, KV heads, head size),
else 32k. Overshooting is
- worse than being conservative: Ollama then moves layers to the CPU and
generation slows to a crawl.
-
-The static prefix of system prompt and core tools is about 4.5k tokens, and
each question with tool
-calls adds another 2k to 4k of history. The panel compacts the history once
the prompt Ollama measured
-for the last request passes half the window, capped at half of 64k even when a
larger window was
-adopted: every token of history is prefill time again after a cache loss (an
idle unload after 30
-minutes, another client, a restart), and 64k at a few hundred tokens per
second is already well over a
-minute. Compacting rewrites the history, which invalidates Ollama's prompt
cache, so the request after
-a compaction prefills the whole prompt again: about 40 seconds for a 21k-token
prompt. The panel
-therefore compacts local history rarely and then thoroughly, down to about a
quarter of the window,
-and prints one line saying what it did and how long the next reply will take
to start. `/compact` does
-the same on demand with the same line; `/context` shows the window, the
compaction point and the last
-measured prompt; the panel title shows the fill as `ctx 34%`.
-
-*Choosing the model.* With no `camel.tui.ai.model` set the panel takes
`llama3.2` when it is installed,
-otherwise the first installed model; run `/model <name>` in the panel or set
*AI Model* in
-*F2 -> Settings* to pin one. The model must support tool calling and should
have at least 14B
-parameters; a mixture-of-experts model such as `qwen3.6:35b-a3b` prefills
several times faster than a
-dense model of similar quality, which is what matters for a tool-heavy prompt.
*F2 -> Run Doctor* shows
-whether Ollama was found, which models are installed and whether they are
large enough.
-
-*Where to look.* The <<_ollama>> tab is the instrument for all of the above:
tokens per second live and
-per request, time to first token with cold starts marked, cache hit, how full
the context window is
-and how it grows per question, GPU and process load, and one line per question
with the requests it
-took, the question being answered right now on top with `working` as its
reason and its time counting
-up, and a footer with the average per question. In the AI panel, `/context`
prints what the next request will cost, `/usage` and *Ctrl+U* the
-session totals per question.
-
-*Remote and containerised Ollama.* Everything above applies to an Ollama on
another host or inside
-`camel infra run ollama` as well, with two differences: the container runs
without GPU acceleration,
-and the live runner state and host load on the Ollama tab need the server on
the same machine.
-
-For the wider picture see the blog posts
-link:/blog/2026/09/camel-local-model-benchmark/[We had a frontier AI coach a
small local model through Camel]
-on what a local model can do with Camel and what was changed to help it, and
-link:/blog/2026/09/camel-genai-observability-jbang/[Observe Your Camel AI
Routes with GenAI OpenTelemetry]
-on observing routes that call Ollama.
-
-=== Using an OpenAI-compatible local server
-
-Set `LLM_API_KEY` and `LLM_BASE_URL` to connect to any OpenAI-compatible server
-(LM Studio, vLLM, llama.cpp, GPT4All, …):
-
-[source,bash]
-----
-export LLM_API_KEY=any-value
-export LLM_BASE_URL=http://localhost:1234
-camel tui
-----
-
-`OPENAI_BASE_URL` is also accepted as an alternative to `LLM_BASE_URL`.
-
-The panel uses the first model the server lists on `/v1/models`. To use
another one, run
-`/model <name>` in the panel (`/model` alone lists what the server offers),
set *AI Model* in
-*F2 -> Settings*, or set `camel.tui.ai.model`. The model must support tool
calling, otherwise the
-panel answers from training data instead of inspecting your integration. When
a request fails, the
-panel shows the HTTP status and the server's error message, for example a
model that the server
-does not host.
+To run the model on your own machine, with Ollama or an OpenAI-compatible
server such as LM Studio, see
+xref:camel-jbang-tui-local-models.adoc[Camel TUI Local Models]: which models
work well, the tool set they get,
+what a question costs and the Ollama tab.
== AI panel slash commands
@@ -228,7 +76,7 @@ cycles backward.
| Show what the next request costs: provider and model, tool set, static
prefix size, history size and the session total, and with Ollama the context
window, the prompt size above which the history is compacted and the last
measured prompt. Useful with local models, where prompt size is time.
| `/compact`
-| Shrink the conversation history sent to the model right away: older tool
results are cut to their first lines and the oldest turns are dropped. With a
hosted provider the panel does this automatically after each answer for all but
the latest turn. With Ollama or another `localhost` provider it waits until the
prompt Ollama measured passes half the context window (see
<<_working_with_a_local_ollama_model>>), then compacts thoroughly, because a
local server can reuse its cached prompt on [...]
+| Shrink the conversation history sent to the model right away: older tool
results are cut to their first lines and the oldest turns are dropped. With a
hosted provider the panel does this automatically after each answer for all but
the latest turn. With Ollama or another `localhost` provider it waits until the
prompt Ollama measured passes half the context window (see
xref:camel-jbang-tui-local-models.adoc#_working_with_a_local_ollama_model[Working
with a local Ollama model]), then comp [...]
| `/retry`
| Send the last question again, starting from a clean turn in the model
history.
@@ -314,7 +162,7 @@ model called and the time spent in them, the number of
requests, the tokens and,
the context window got (for example `61.1s · 25 tool calls, limit reached
(2.5s in tools) · 26 requests
· 231.3k tokens · ctx 18%`). The panel allows 25 tool calls per question. When
the model reaches that
the byline says `limit reached`, the panel asks the model to answer from what
it has found so far, and
-the <<_ollama>> tab shows the question's request count in red with `limit` as
its reason (the count is
+the xref:camel-jbang-tui-local-models.adoc#_ollama_tab[Ollama tab] shows the
question's request count in red with `limit` as its reason (the count is
yellow from 10 requests, and the total time yellow from 30 seconds and orange
from a minute). A question
that runs into the limit is usually a tool that makes the model guess, not a
weak model: the AI log
shows which one.
@@ -353,45 +201,3 @@ on a 10k-token prompt means the cache was hit, while a
prefill of several second
prompt was processed again. For OpenAI, Anthropic and Gemini it prints the
number of `cached` input
tokens the provider reported. Use it to check that follow-up questions are
cheap before blaming the
model for being slow.
-
-== Ollama
-
-The Ollama tab (under More, in the *AI* group) shows how the model served by a
local
-https://ollama.com[Ollama] is performing, in the spirit of an LLM dashboard.
It works with or without a
-running integration: it finds Ollama at `localhost:11434`, at the address of
`camel infra run ollama`,
-or at the endpoint the AI panel (*F8*) is using. The tab is listed only while
an Ollama server answers;
-the TUI checks every ten seconds, so it appears shortly after `ollama serve`
starts.
-
-* *Model* -- the loaded model with its family, parameters, quantization,
layers, experts (and how many
- are active per token for a mixture-of-experts model), how much of it sits in
GPU memory, the
- allocated context length and when Ollama will unload it. With no model
loaded, the installed models
- are listed instead.
-* *Throughput* -- decode and prefill tokens per second, *live* while the model
is generating and
- otherwise from the last request; time to first token and load time (a load
of a second or more
- is a cold start); session averages and a sparkline of the decode rate.
-* *Context* -- how full the context window is, the share of the prompt served
from Ollama's cache,
- whether the model is working or idle, the speculative decoding method in
use, and a per-turn trend
- of how much of the window each AI panel prompt filled, with the session peak
and the compactions seen.
-* *Host* -- GPU utilization and memory (Apple silicon through `ioreg`, NVIDIA
through `nvidia-smi`),
- and CPU and memory of the Ollama server and its model runner.
-* *Requests* -- one line per question asked in the AI panel (a question with
tool calls costs one
- request per step; *Enter* unfolds the steps) with the question text, prompt
and generated tokens,
- cache hit, the share of the context window reached, prefill and decode
tokens per second, time to
- first token, the time you waited and the stop reason. Calls made by Camel
routes are listed as
- their own lines.
-
-Two kinds of requests appear in the log. Questions asked in the AI panel with
Ollama as the provider
-come with the timings Ollama reports for each request (prompt evaluation,
generation, model load,
-total). Calls made by Camel routes through `camel-langchain4j-chat`,
`camel-openai` or
-`camel-spring-ai-chat` appear when the integration runs with GenAI
observability (`--observe`, or
-`--dep=camel:ai-observability`), tagged with the route id; Camel records
tokens and duration for those,
-not the phase split.
-
-The live figures, the context panel and the host panel need the model runner
on the same machine:
-Ollama starts a `llama-server` process per loaded model and the tab reads its
slot state a few times a
-second. Against a remote or containerised Ollama the tab keeps the model,
per-request and session data
-and says which panels are unavailable.
-
-Press *r* to reset the request log and the session totals, *F5* to refresh
immediately. The same data
-is available to AI agents through the `tui_get_ollama` MCP tool. For what the
figures mean for the AI
-panel and which knobs to turn, see <<_working_with_a_local_ollama_model>>.
diff --git
a/docs/user-manual/modules/ROOT/pages/camel-jbang-tui-getting-started.adoc
b/docs/user-manual/modules/ROOT/pages/camel-jbang-tui-getting-started.adoc
new file mode 100644
index 000000000000..f7f854f0f5e9
--- /dev/null
+++ b/docs/user-manual/modules/ROOT/pages/camel-jbang-tui-getting-started.adoc
@@ -0,0 +1,124 @@
+= Camel TUI Getting Started
+
+There are several ways to get going with the TUI: point it at integrations
that already run, run one of the
+built-in examples, open a project directory, or connect an existing Spring
Boot or Quarkus application.
+
+See xref:camel-jbang-tui.adoc[Camel TUI] for the tabs, the keyboard shortcuts
and the other pages.
+
+== Option 1: Your Own Route
+
+Start a Camel integration in one terminal:
+
+[source,bash]
+----
+camel run my-route.yaml
+----
+
+Open the TUI in another terminal:
+
+[source,bash]
+----
+camel tui
+----
+
+The TUI auto-discovers every running Camel integration on your machine -- no
configuration needed.
+
+== Option 2: Built-in Examples
+
+Don't have a route yet? The TUI ships with a catalog of ready-to-run examples.
+Open the TUI and press *F2*, then select _Run an example_:
+
+[source,bash]
+----
+camel tui
+----
+
+. Press *F2* to open the actions menu
+. Select *Run an example...*
+. Browse the example catalog -- type to filter by name
+. Press *Enter* to launch the selected example
+
+Before launching, a run options form lets you choose the *runtime*: Camel Main
(standalone),
+Spring Boot, Quarkus, or JBang. This makes it easy to try any example on all
three runtimes without
+changing a single line of code. The first three run the example in a separate
JVM that only
+contains the dependencies of the example (like a production deployment), while
JBang runs it
+in-process in the Camel CLI JVM, which starts faster but has the CLI on the
classpath as well. You can also set the integration name, toggle dev mode, and
+add extra dependencies. When the port the app will listen on (the one you set,
else 8080) is taken already, the form
+says by whom, so a second app with an HTTP server does not fail to start.
+
+The example starts running in the background. The TUI auto-selects it as soon
as it appears.
+From there you can explore tabs, watch messages flow, inspect the route
diagram, and experiment.
+
+If an example requires infrastructure (like Kafka or a database), the TUI
automatically starts the
+required Docker containers before launching the example. A notification in the
footer shows the
+progress.
+
+The same runtime selector is available in *Run from folder...* (F2 menu),
which lets you point
+the TUI at a local directory containing your routes. When a `pom.xml` is
present, the runtime
+is auto-detected and locked to match your project.
+
+TIP: Press *F1* or *?* on any screen for context-sensitive help.
+Keyboard shortcuts are always shown in the footer bar.
+
+== Option 3: Open a Project Directory
+
+You can point the TUI directly at a project directory:
+
+[source,bash]
+----
+camel tui .
+camel tui /path/to/my-project
+----
+
+The TUI opens the directory in the Source tab so you can browse the project
files immediately.
+When a `pom.xml` is present, the runtime is auto-detected (Spring Boot,
Quarkus, or Camel Main).
+Press *F10* to run the project -- Maven projects are run with `camel run
pom.xml`, which launches
+them with the appropriate goal (`spring-boot:run`, `quarkus:dev`, or
`camel:run`), and plain
+directories are run with `camel run`. Either way the application logs to
`~/.camel` so the Log tab
+shows its logs.
+
+Until it runs, the project is listed as *Stopped*: the Overview says what it
is (runtime, Maven or route files, its
+folder) and the Source pane lists the routes it found and where they are, so
*g* opens one. While it runs, the app
+takes its place under the name it gives itself, with the project's name next
to it; when the run ends, the project is
+back as Stopped and still selected.
+
+This is a quick way to explore and run any Camel project without starting it
separately first.
+
+image::jbang/camel-tui-main-quarkus-source.png[A Quarkus project in the Source
tab: its Java route and the bean of the line]
+
+== Connecting Existing Spring Boot and Quarkus Applications
+
+The TUI auto-discovers integrations started with `camel run`. To monitor and
control
+existing Spring Boot or Quarkus applications, add the `camel-cli-connector`
dependency
+to your project. This is a lightweight runtime adapter that lets the TUI (and
the Camel CLI)
+communicate with your application -- all TUI features work the same way
regardless of runtime.
+
+Spring Boot:
+
+[source,xml]
+----
+<dependency>
+ <groupId>org.apache.camel.springboot</groupId>
+ <artifactId>camel-cli-connector-starter</artifactId>
+</dependency>
+----
+
+Quarkus:
+
+[source,xml]
+----
+<dependency>
+ <groupId>org.apache.camel.quarkus</groupId>
+ <artifactId>camel-quarkus-cli-connector</artifactId>
+</dependency>
+----
+
+Once added, start your application normally and the TUI will discover it
automatically.
+No additional configuration is needed -- the connector auto-detects on the
classpath and
+registers the application with the local Camel CLI.
+
+TIP: When you open a Spring Boot, Quarkus or Camel Main project via *F2 > Open
Project* and run
+it with *F10*, the CLI connector dependency is automatically injected if it's
not already in your
+`pom.xml`. This means the TUI can monitor the application without modifying
your project.
+
+See xref:camel-jbang-managing.adoc[Managing Integrations] for more details.
diff --git
a/docs/user-manual/modules/ROOT/pages/camel-jbang-tui-local-models.adoc
b/docs/user-manual/modules/ROOT/pages/camel-jbang-tui-local-models.adoc
new file mode 100644
index 000000000000..aacdf2def9bf
--- /dev/null
+++ b/docs/user-manual/modules/ROOT/pages/camel-jbang-tui-local-models.adoc
@@ -0,0 +1,213 @@
+= Camel TUI Local Models
+
+The xref:camel-jbang-tui-ai.adoc[AI panel] (*F8*) works with a model running
on your own machine, through
+https://ollama.com[Ollama] or any OpenAI-compatible server, so no API key is
needed and nothing has to leave
+your machine. A local model spends its time differently from a hosted one, and
the TUI shows you where.
+
+See xref:camel-jbang-tui.adoc[Camel TUI] for getting started and the other
pages.
+
+== Key Features
+
+* <<_using_ollama_local_no_api_key,*Ollama*>> -- auto-detected at
`localhost:11434`, with the models that work well
+* <<_using_an_openai_compatible_local_server,*OpenAI-compatible servers*>> --
LM Studio, vLLM, llama.cpp and others
+* <<_tool_set_for_local_models,*A smaller tool set*>> -- local models get the
core tools, which roughly halves the prompt
+* <<_working_with_a_local_ollama_model,*What a question costs*>> -- load,
prefill and decode, the prompt cache and the context window
+* <<_ollama_tab,*Ollama tab*>> -- the loaded model, tokens per second, cache
hit and context fill, live
+
+== Using Ollama (local, no API key)
+
+Install Ollama natively for best performance — the native binary uses GPU
acceleration
+(Metal on macOS, CUDA/ROCm on Linux):
+
+[source,bash]
+----
+# macOS
+brew install ollama
+
+# Linux
+curl -fsSL https://ollama.com/install.sh | sh
+
+# Pull a model — then open the TUI and press F8
+ollama pull qwen3.6:35b-a3b
+camel tui
+----
+
+Ollama at `localhost:11434` is auto-detected. No configuration needed.
+
+IMPORTANT: The F8 AI panel works by invoking built-in tools to inspect your
running Camel process.
+Models smaller than ~14B do not reliably call tools and answer from training
knowledge instead.
+Use at least a 14B model. Prefer a mixture-of-experts model such as
`qwen3.6:35b-a3b`: with only
+3B parameters active per token it processes the tool-heavy prompt many times
faster than a dense
+27B/32B model, so answers start in seconds instead of a minute.
+
+*Models that work well* (tool-calling capable, ≥14B, default Q4_K_M
quantization):
+
+[options="header"]
+|===
+| Model | RAM | Notes
+| `qwen3.6:35b-a3b` | ~23 GB | Recommended: fastest prompt processing, needs
32 GB+
+| `qwen2.5:14b` | ~9 GB | Minimum for 16 GB machines
+| `qwen3.6:27b` | ~18 GB | Strong dense model, several times slower prompt
processing
+| `qwen2.5:32b` | ~20 GB | Good quality, slow prompt processing
+| `hermes3:70b` | ~43 GB | Excellent tool calling, needs 64 GB+
+| `llama3.3:70b` | ~43 GB | Best open model, needs 64 GB+
+|===
+
+NOTE: `camel infra run ollama` runs Ollama in Docker and bypasses GPU
acceleration,
+making inference significantly slower. Native install is preferred for
development.
+
+NOTE: On Apple Silicon, use the default (GGUF) tags rather than the `-mlx`
tags. The Ollama MLX
+engine cannot yet reuse the cached prompt for Qwen 3.x models, so every
question re-processes the
+whole prompt, while the default engine reuses it and only processes what is
new.
+
+== Using an OpenAI-compatible local server
+
+Set `LLM_API_KEY` and `LLM_BASE_URL` to connect to any OpenAI-compatible server
+(LM Studio, vLLM, llama.cpp, GPT4All, …):
+
+[source,bash]
+----
+export LLM_API_KEY=any-value
+export LLM_BASE_URL=http://localhost:1234
+camel tui
+----
+
+`OPENAI_BASE_URL` is also accepted as an alternative to `LLM_BASE_URL`.
+
+The panel uses the first model the server lists on `/v1/models`. To use
another one, run
+`/model <name>` in the panel (`/model` alone lists what the server offers),
set *AI Model* in
+*F2 -> Settings*, or set `camel.tui.ai.model`. The model must support tool
calling, otherwise the
+panel answers from training data instead of inspecting your integration. When
a request fails, the
+panel shows the HTTP status and the server's error message, for example a
model that the server
+does not host.
+
+== Tool set for local models
+
+Every question sends the definitions of the `tui_*` tools the model may call,
and a local model
+pays for each of them in prompt-processing time. The panel therefore sends
only the core set of
+tools (state, tables, logs, errors, diagrams, topology, processor details,
catalog docs, traces,
+spans, route control, sending messages, source files, infra services,
navigation, log level and
+filters) to Ollama
+and to any provider on `localhost`, which roughly halves the prompt. Hosted
providers get every
+tool, including the drawing, animation and automation tools. Use `/tools full`
in the panel to send
+all tools to a local model too, `/tools core` to trim the set for a hosted
one, pick *AI Tools* in
+*F2 -> Settings*, or set `camel.tui.ai.tools` in `.camel-cli.properties`.
Ollama requests also ask the server to keep the
+model loaded for 30 minutes and for a context window of 32k or 64k (see
<<_working_with_a_local_ollama_model>>;
+`OLLAMA_CONTEXT_LENGTH` overrides it), so follow-up questions reuse the cached
prompt instead of reloading
+the model.
+
+== Working with a local Ollama model
+
+A local model is not a slower version of a hosted one; it spends its time
differently, and the TUI
+shows you where. This section explains what a question costs with Ollama and
which knobs matter,
+with figures measured on an Apple M4 Pro (64 GB) running `qwen3.6:35b-a3b`.
+
+*How a question is spent.* Ollama answers in three phases, and every timing
the TUI shows maps to one
+of them:
+
+1. *Load*: if the model is not in memory (first question, the keep-alive
expired, or a request asked for
+ a different context size) Ollama starts a runner and loads the weights: 10
to 20 seconds for a 22 GB
+ model. The TUI asks Ollama to keep the model loaded for 30 minutes after
each request.
+2. *Prefill*: the prompt (system prompt, tool definitions, conversation
history, your question) is
+ processed in one batch, at roughly 600 to 700 tokens per second when
nothing is cached.
+3. *Decode*: the answer is generated token by token, at 50 to 60 tokens per
second for this model.
+
+The wait before anything appears, the time to first token, is load plus
prefill. A cold first question
+therefore takes 20 seconds before the first word; a warm one under a second.
+
+*One question, many requests.* The panel answers by calling the `tui_*` tools,
and every tool call the
+model makes costs a new request that sends the whole prompt again. A simple
question can take 3 to 13
+requests. This is affordable only because Ollama caches the prompt prefix: the
system prompt, the tool
+definitions and the history are identical from one step to the next, so each
step prefills only the new
+tokens and takes about half a second. Across a session the cache hit is
typically above 90%. The
+Ollama tab shows the request count per question as `×N` and the cache hit per
question; a question
+with a high count and a short answer is the model exploring, which a smaller,
sharper tool set reduces
+(see <<_tool_set_for_local_models>>).
+
+*The context window.* The TUI decides the window it asks Ollama for
(`num_ctx`) once per model:
+
+1. `OLLAMA_CONTEXT_LENGTH` in the environment wins when set.
+2. If the model is already loaded, its window is adopted (raised to 32k when
smaller), so the TUI never
+ makes Ollama reload a model another client is using. A request with a
different `num_ctx` costs a
+ cold start and throws away everyone's prompt cache.
+3. Otherwise 64k when the model's weights plus the KV cache of a 64k window
fit the machine's memory,
+ computed from what `ollama show` reports (layers, KV heads, head size),
else 32k. Overshooting is
+ worse than being conservative: Ollama then moves layers to the CPU and
generation slows to a crawl.
+
+The static prefix of system prompt and core tools is about 4.5k tokens, and
each question with tool
+calls adds another 2k to 4k of history. The panel compacts the history once
the prompt Ollama measured
+for the last request passes half the window, capped at half of 64k even when a
larger window was
+adopted: every token of history is prefill time again after a cache loss (an
idle unload after 30
+minutes, another client, a restart), and 64k at a few hundred tokens per
second is already well over a
+minute. Compacting rewrites the history, which invalidates Ollama's prompt
cache, so the request after
+a compaction prefills the whole prompt again: about 40 seconds for a 21k-token
prompt. The panel
+therefore compacts local history rarely and then thoroughly, down to about a
quarter of the window,
+and prints one line saying what it did and how long the next reply will take
to start. `/compact` does
+the same on demand with the same line; `/context` shows the window, the
compaction point and the last
+measured prompt; the panel title shows the fill as `ctx 34%`.
+
+*Choosing the model.* With no `camel.tui.ai.model` set the panel takes
`llama3.2` when it is installed,
+otherwise the first installed model; run `/model <name>` in the panel or set
*AI Model* in
+*F2 -> Settings* to pin one. The model must support tool calling and should
have at least 14B
+parameters; a mixture-of-experts model such as `qwen3.6:35b-a3b` prefills
several times faster than a
+dense model of similar quality, which is what matters for a tool-heavy prompt.
*F2 -> Run Doctor* shows
+whether Ollama was found, which models are installed and whether they are
large enough.
+
+*Where to look.* The <<_ollama_tab,Ollama tab>> is the instrument for all of
the above: tokens per second live and
+per request, time to first token with cold starts marked, cache hit, how full
the context window is
+and how it grows per question, GPU and process load, and one line per question
with the requests it
+took, the question being answered right now on top with `working` as its
reason and its time counting
+up, and a footer with the average per question. In the AI panel, `/context`
prints what the next request will cost, `/usage` and *Ctrl+U* the
+session totals per question.
+
+*Remote and containerised Ollama.* Everything above applies to an Ollama on
another host or inside
+`camel infra run ollama` as well, with two differences: the container runs
without GPU acceleration,
+and the live runner state and host load on the Ollama tab need the server on
the same machine.
+
+For the wider picture see the blog posts
+link:/blog/2026/09/camel-local-model-benchmark/[We had a frontier AI coach a
small local model through Camel]
+on what a local model can do with Camel and what was changed to help it, and
+link:/blog/2026/09/camel-genai-observability-jbang/[Observe Your Camel AI
Routes with GenAI OpenTelemetry]
+on observing routes that call Ollama.
+
+== Ollama tab
+
+The Ollama tab (under More, in the *AI* group) shows how the model served by a
local
+https://ollama.com[Ollama] is performing, in the spirit of an LLM dashboard.
It works with or without a
+running integration: it finds Ollama at `localhost:11434`, at the address of
`camel infra run ollama`,
+or at the endpoint the AI panel (*F8*) is using. The tab is listed only while
an Ollama server answers;
+the TUI checks every ten seconds, so it appears shortly after `ollama serve`
starts.
+
+* *Model* -- the loaded model with its family, parameters, quantization,
layers, experts (and how many
+ are active per token for a mixture-of-experts model), how much of it sits in
GPU memory, the
+ allocated context length and when Ollama will unload it. With no model
loaded, the installed models
+ are listed instead.
+* *Throughput* -- decode and prefill tokens per second, *live* while the model
is generating and
+ otherwise from the last request; time to first token and load time (a load
of a second or more
+ is a cold start); session averages and a sparkline of the decode rate.
+* *Context* -- how full the context window is, the share of the prompt served
from Ollama's cache,
+ whether the model is working or idle, the speculative decoding method in
use, and a per-turn trend
+ of how much of the window each AI panel prompt filled, with the session peak
and the compactions seen.
+* *Host* -- GPU utilization and memory (Apple silicon through `ioreg`, NVIDIA
through `nvidia-smi`),
+ and CPU and memory of the Ollama server and its model runner.
+* *Requests* -- one line per question asked in the AI panel (a question with
tool calls costs one
+ request per step; *Enter* unfolds the steps) with the question text, prompt
and generated tokens,
+ cache hit, the share of the context window reached, prefill and decode
tokens per second, time to
+ first token, the time you waited and the stop reason. Calls made by Camel
routes are listed as
+ their own lines.
+
+Two kinds of requests appear in the log. Questions asked in the AI panel with
Ollama as the provider
+come with the timings Ollama reports for each request (prompt evaluation,
generation, model load,
+total). Calls made by Camel routes through `camel-langchain4j-chat`,
`camel-openai` or
+`camel-spring-ai-chat` appear when the integration runs with GenAI
observability (`--observe`, or
+`--dep=camel:ai-observability`), tagged with the route id; Camel records
tokens and duration for those,
+not the phase split.
+
+The live figures, the context panel and the host panel need the model runner
on the same machine:
+Ollama starts a `llama-server` process per loaded model and the tab reads its
slot state a few times a
+second. Against a remote or containerised Ollama the tab keeps the model,
per-request and session data
+and says which panels are unavailable.
+
+Press *r* to reset the request log and the session totals, *F5* to refresh
immediately. The same data
+is available to AI agents through the `tui_get_ollama` MCP tool. For what the
figures mean for the AI
+panel and which knobs to turn, see <<_working_with_a_local_ollama_model>>.
diff --git a/docs/user-manual/modules/ROOT/pages/camel-jbang-tui.adoc
b/docs/user-manual/modules/ROOT/pages/camel-jbang-tui.adoc
index 2bb536b6a25d..78f90dddee3e 100644
--- a/docs/user-manual/modules/ROOT/pages/camel-jbang-tui.adoc
+++ b/docs/user-manual/modules/ROOT/pages/camel-jbang-tui.adoc
@@ -14,7 +14,7 @@ image::jbang/camel-tui-overview.png[TUI Overview showing
multiple routes]
== Key Features
* *Run anything* -- your own routes, the built-in examples, or an existing
Spring Boot or Quarkus project
- (<<_getting_started,Getting started>>).
+ (xref:camel-jbang-tui-getting-started.adoc[Getting started]).
* *Write routes with help* -- a xref:camel-jbang-tui-source-editor.adoc[source
editor] that checks YAML, Java and XML
routes as you type, completes endpoint options, fixes problems and shows the
live run data of each line.
* *See the integration* -- the xref:camel-jbang-tui-diagram.adoc[diagram] from
the architecture down to one route,
@@ -28,126 +28,23 @@ image::jbang/camel-tui-overview.png[TUI Overview showing
multiple routes]
== Getting Started
-You can start using the TUI in two ways: with your own route, or by running
one of the built-in examples.
-
-=== Option 1: Your Own Route
-
-Start a Camel integration in one terminal:
-
[source,bash]
----
-camel run my-route.yaml
-----
-
-Open the TUI in another terminal:
-
-[source,bash]
-----
-camel tui
-----
-
-The TUI auto-discovers every running Camel integration on your machine -- no
configuration needed.
-
-=== Option 2: Built-in Examples
+camel run my-route.yaml # in one terminal
+camel tui # in another: auto-discovers every running
integration
-Don't have a route yet? The TUI ships with a catalog of ready-to-run examples.
-Open the TUI and press *F2*, then select _Run an example_:
-
-[source,bash]
-----
-camel tui
+camel tui . # or open a project directory and press F10 to run it
----
-. Press *F2* to open the actions menu
-. Select *Run an example...*
-. Browse the example catalog -- type to filter by name
-. Press *Enter* to launch the selected example
-
-Before launching, a run options form lets you choose the *runtime*: Camel Main
(standalone),
-Spring Boot, Quarkus, or JBang. This makes it easy to try any example on all
three runtimes without
-changing a single line of code. The first three run the example in a separate
JVM that only
-contains the dependencies of the example (like a production deployment), while
JBang runs it
-in-process in the Camel CLI JVM, which starts faster but has the CLI on the
classpath as well. You can also set the integration name, toggle dev mode, and
-add extra dependencies. When the port the app will listen on (the one you set,
else 8080) is taken already, the form
-says by whom, so a second app with an HTTP server does not fail to start.
-
-The example starts running in the background. The TUI auto-selects it as soon
as it appears.
-From there you can explore tabs, watch messages flow, inspect the route
diagram, and experiment.
-
-If an example requires infrastructure (like Kafka or a database), the TUI
automatically starts the
-required Docker containers before launching the example. A notification in the
footer shows the
-progress.
+No route yet? Press *F2* and pick _Run an example..._ to run one of the
built-in examples on Camel Main,
+Spring Boot, Quarkus or JBang.
-The same runtime selector is available in *Run from folder...* (F2 menu),
which lets you point
-the TUI at a local directory containing your routes. When a `pom.xml` is
present, the runtime
-is auto-detected and locked to match your project.
+See xref:camel-jbang-tui-getting-started.adoc[Camel TUI Getting Started] for
each way to start, and for
+connecting existing Spring Boot and Quarkus applications.
TIP: Press *F1* or *?* on any screen for context-sensitive help.
Keyboard shortcuts are always shown in the footer bar.
-=== Option 3: Open a Project Directory
-
-You can point the TUI directly at a project directory:
-
-[source,bash]
-----
-camel tui .
-camel tui /path/to/my-project
-----
-
-The TUI opens the directory in the Source tab so you can browse the project
files immediately.
-When a `pom.xml` is present, the runtime is auto-detected (Spring Boot,
Quarkus, or Camel Main).
-Press *F10* to run the project -- Maven projects are run with `camel run
pom.xml`, which launches
-them with the appropriate goal (`spring-boot:run`, `quarkus:dev`, or
`camel:run`), and plain
-directories are run with `camel run`. Either way the application logs to
`~/.camel` so the Log tab
-shows its logs.
-
-Until it runs, the project is listed as *Stopped*: the Overview says what it
is (runtime, Maven or route files, its
-folder) and the Source pane lists the routes it found and where they are, so
*g* opens one. While it runs, the app
-takes its place under the name it gives itself, with the project's name next
to it; when the run ends, the project is
-back as Stopped and still selected.
-
-This is a quick way to explore and run any Camel project without starting it
separately first.
-
-image::jbang/camel-tui-main-quarkus-source.png[A Quarkus project in the Source
tab: its Java route and the bean of the line]
-
-=== Connecting Existing Spring Boot and Quarkus Applications
-
-The TUI auto-discovers integrations started with `camel run`. To monitor and
control
-existing Spring Boot or Quarkus applications, add the `camel-cli-connector`
dependency
-to your project. This is a lightweight runtime adapter that lets the TUI (and
the Camel CLI)
-communicate with your application -- all TUI features work the same way
regardless of runtime.
-
-Spring Boot:
-
-[source,xml]
-----
-<dependency>
- <groupId>org.apache.camel.springboot</groupId>
- <artifactId>camel-cli-connector-starter</artifactId>
-</dependency>
-----
-
-Quarkus:
-
-[source,xml]
-----
-<dependency>
- <groupId>org.apache.camel.quarkus</groupId>
- <artifactId>camel-quarkus-cli-connector</artifactId>
-</dependency>
-----
-
-Once added, start your application normally and the TUI will discover it
automatically.
-No additional configuration is needed -- the connector auto-detects on the
classpath and
-registers the application with the local Camel CLI.
-
-TIP: When you open a Spring Boot, Quarkus or Camel Main project via *F2 > Open
Project* and run
-it with *F10*, the CLI connector dependency is automatically injected if it's
not already in your
-`pom.xml`. This means the TUI can monitor the application without modifying
your project.
-
-See xref:camel-jbang-managing.adoc[Managing Integrations] for more details.
-
== Tabs Overview
The TUI organizes information into tabs. Press number keys *1* through *0* to
jump directly
@@ -190,30 +87,22 @@ Two panels can be opened on top of any tab: *F6* opens an
xref:camel-jbang-tui-s
for running `camel` commands, and *F8* opens the
xref:camel-jbang-tui-ai.adoc[AI panel] for asking
questions about the running integrations.
-== Source Code Browser
-
-The Source tab (Tab 2) is a file explorer and editor for your project code
that knows Camel. It checks your YAML,
-Java and XML routes as you type, completes endpoint options from the Camel
catalog, explains the line the cursor is on,
-fixes problems for you (or asks the AI to), and shows the exchanges and
failures of each line of the running
-integration next to the code.
-
-image::jbang/camel-tui-source-live-run-data.png[Live run data in the source
editor]
-
-See xref:camel-jbang-tui-source-editor.adoc[Camel TUI Source Editor] for all
its features and keys.
-
== More about the TUI
[cols="1,3",options="header"]
|===
| Page | What it covers
+| xref:camel-jbang-tui-getting-started.adoc[Getting Started] | Your own
routes, the built-in examples, opening a project,
+connecting Spring Boot and Quarkus applications
| xref:camel-jbang-tui-source-editor.adoc[Source Editor] | Reading and writing
routes: checks as you type, quick fixes,
fix with AI, completion, quick documentation, navigation, live run data
| xref:camel-jbang-tui-diagram.adoc[Diagram] | The architecture, topology and
route views, external endpoints, metrics
| xref:camel-jbang-tui-observe.adoc[Observing Integrations] | Activity,
message history, errors, spans, process, HTTP
probe, CVE audit, Kafka, SQL, memory leaks, JFR, catalog
-| xref:camel-jbang-tui-ai.adoc[AI Panel] | AI providers and local models,
slash commands, project overview, AI log, the
-Ollama tab
+| xref:camel-jbang-tui-ai.adoc[AI Panel] | AI providers, slash commands,
project overview, AI log
+| xref:camel-jbang-tui-local-models.adoc[Local Models] | Ollama and
OpenAI-compatible servers, the tool set for local
+models, what a question costs, the Ollama tab
| xref:camel-jbang-tui-ai-agents.adoc[AI Agents] | MCP and ACP: what an AI
agent can see and do, edits you confirm
| xref:camel-jbang-tui-settings.adoc[Actions, Settings and Themes] | The
actions menu, embedded shell, themes, settings,
browser access, recording demos
@@ -276,7 +165,7 @@ See the
xref:camel-jbang-tui-source-editor.adoc#_keyboard_shortcuts[keyboard sho
| `100`
| `--theme`
-| Color theme for this session (e.g., `dark`, `tokyo-night`, `dracula`). See
<<Theme>> for the full list of 21 themes. Overrides the persisted
`camel.tui.theme` preference when set.
+| Color theme for this session (e.g., `dark`, `tokyo-night`, `dracula`). See
xref:camel-jbang-tui-settings.adoc#_theme[Theme] for the full list of 21
themes. Overrides the persisted `camel.tui.theme` preference when set.
|
| `--record`