This is an automated email from the ASF dual-hosted git repository.

davsclaus pushed a commit to branch main
in repository https://gitbox.apache.org/repos/asf/camel.git


The following commit(s) were added to refs/heads/main by this push:
     new 4062f82648b4 chore: docs - camel-jbang local models and getting 
started get their own pages (#27282)
4062f82648b4 is described below

commit 4062f82648b484736219c804141f668a6ec66bcb
Author: Claus Ibsen <[email protected]>
AuthorDate: Fri Oct 2 15:06:45 2026 +0200

    chore: docs - camel-jbang local models and getting started get their own 
pages (#27282)
    
    Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
---
 docs/user-manual/modules/ROOT/nav.adoc             |   2 +
 .../modules/ROOT/pages/camel-jbang-ai.adoc         |   2 +-
 .../modules/ROOT/pages/camel-jbang-tui-ai.adoc     | 208 +-------------------
 .../pages/camel-jbang-tui-getting-started.adoc     | 124 ++++++++++++
 .../ROOT/pages/camel-jbang-tui-local-models.adoc   | 213 +++++++++++++++++++++
 .../modules/ROOT/pages/camel-jbang-tui.adoc        | 139 ++------------
 6 files changed, 361 insertions(+), 327 deletions(-)

diff --git a/docs/user-manual/modules/ROOT/nav.adoc 
b/docs/user-manual/modules/ROOT/nav.adoc
index f3a51c5d0dc2..08c67e8e6c96 100644
--- a/docs/user-manual/modules/ROOT/nav.adoc
+++ b/docs/user-manual/modules/ROOT/nav.adoc
@@ -11,10 +11,12 @@
 *** xref:camel-jbang-kubernetes.adoc[Camel Kubernetes Plugin]
 *** xref:camel-jbang-test.adoc[Camel Testing Plugin]
 *** xref:camel-jbang-tui.adoc[Camel TUI]
+**** xref:camel-jbang-tui-getting-started.adoc[Camel TUI Getting Started]
 **** xref:camel-jbang-tui-source-editor.adoc[Camel TUI Source Editor]
 **** xref:camel-jbang-tui-diagram.adoc[Camel TUI Diagram]
 **** xref:camel-jbang-tui-observe.adoc[Camel TUI Observing Integrations]
 **** xref:camel-jbang-tui-ai.adoc[Camel TUI AI Panel]
+**** xref:camel-jbang-tui-local-models.adoc[Camel TUI Local Models]
 **** xref:camel-jbang-tui-ai-agents.adoc[Camel TUI and AI Agents]
 **** xref:camel-jbang-tui-settings.adoc[Camel TUI Actions, Settings and Themes]
 *** xref:camel-jbang-mcp.adoc[Camel MCP Server]
diff --git a/docs/user-manual/modules/ROOT/pages/camel-jbang-ai.adoc 
b/docs/user-manual/modules/ROOT/pages/camel-jbang-ai.adoc
index 2526cfb71d1c..8ce276952788 100644
--- a/docs/user-manual/modules/ROOT/pages/camel-jbang-ai.adoc
+++ b/docs/user-manual/modules/ROOT/pages/camel-jbang-ai.adoc
@@ -276,7 +276,7 @@ of reloading the model, and for a context window 
(`num_ctx`) chosen once per mod
 The tool-calling prompt alone is about 4.5k tokens, which is why 32k is the 
floor: Ollama's own default
 on machines with less than 24 GB is 4k and would truncate it. Keep all your 
Ollama clients on the same
 value; if you set `OLLAMA_CONTEXT_LENGTH` for the CLI, set it for `ollama 
serve` too. The TUI's Ollama
-tab shows the window in use, and 
xref:camel-jbang-tui-ai.adoc#_working_with_a_local_ollama_model[Working
+tab shows the window in use, and 
xref:camel-jbang-tui-local-models.adoc#_working_with_a_local_ollama_model[Working
 with a local Ollama model] explains how the AI panel manages its history 
inside it.
 
 === Using an OpenAI-compatible local server
diff --git a/docs/user-manual/modules/ROOT/pages/camel-jbang-tui-ai.adoc 
b/docs/user-manual/modules/ROOT/pages/camel-jbang-tui-ai.adoc
index 2a1ac2268d8b..33eb1244b06e 100644
--- a/docs/user-manual/modules/ROOT/pages/camel-jbang-tui-ai.adoc
+++ b/docs/user-manual/modules/ROOT/pages/camel-jbang-tui-ai.adoc
@@ -2,7 +2,7 @@
 
 The AI panel (*F8*) lets you ask about your integrations in plain language: 
the AI sees what the TUI sees, reads the
 sources and the catalog, and can change files with your confirmation. It works 
with hosted models and with local
-models through Ollama, so nothing has to leave your machine.
+models through Ollama, so nothing has to leave your machine (see 
xref:camel-jbang-tui-local-models.adoc[Local Models]).
 
 See xref:camel-jbang-tui.adoc[Camel TUI] for getting started and the other 
pages.
 
@@ -12,7 +12,7 @@ See xref:camel-jbang-tui.adoc[Camel TUI] for getting started 
and the other pages
 * <<_ai_panel_slash_commands,*Slash commands*>> -- run examples, change 
settings and steer the AI from the panel
 * <<_ai_project_overview,*Project overview*>> -- the AI explains the project 
once and the diagram shows it as capabilities
 * <<_ai_log,*AI log*>> -- what was asked, which tools ran and what they 
returned
-* <<_ollama,*Ollama tab*>> -- the local models, their context window and what 
is loaded
+* xref:camel-jbang-tui-local-models.adoc[*Local models*] -- Ollama or an 
OpenAI-compatible server on your own machine
 
 == Choosing an AI provider
 
@@ -39,161 +39,9 @@ via *F2* -> _AI & MCP_ -> _Setup AI_. Use *F2* -> 
_Settings_ to pin a provider,
 
 image::jbang/camel-tui-ai-panel-answer.png[A local Ollama model explains why 
an order failed]
 
-=== Using Ollama (local, no API key)
-
-Install Ollama natively for best performance — the native binary uses GPU 
acceleration
-(Metal on macOS, CUDA/ROCm on Linux):
-
-[source,bash]
-----
-# macOS
-brew install ollama
-
-# Linux
-curl -fsSL https://ollama.com/install.sh | sh
-
-# Pull a model — then open the TUI and press F8
-ollama pull qwen3.6:35b-a3b
-camel tui
-----
-
-Ollama at `localhost:11434` is auto-detected. No configuration needed.
-
-IMPORTANT: The F8 AI panel works by invoking built-in tools to inspect your 
running Camel process.
-Models smaller than ~14B do not reliably call tools and answer from training 
knowledge instead.
-Use at least a 14B model. Prefer a mixture-of-experts model such as 
`qwen3.6:35b-a3b`: with only
-3B parameters active per token it processes the tool-heavy prompt many times 
faster than a dense
-27B/32B model, so answers start in seconds instead of a minute.
-
-*Models that work well* (tool-calling capable, ≥14B, default Q4_K_M 
quantization):
-
-[options="header"]
-|===
-| Model | RAM | Notes
-| `qwen3.6:35b-a3b` | ~23 GB | Recommended: fastest prompt processing, needs 
32 GB+
-| `qwen2.5:14b` | ~9 GB | Minimum for 16 GB machines
-| `qwen3.6:27b` | ~18 GB | Strong dense model, several times slower prompt 
processing
-| `qwen2.5:32b` | ~20 GB | Good quality, slow prompt processing
-| `hermes3:70b` | ~43 GB | Excellent tool calling, needs 64 GB+
-| `llama3.3:70b` | ~43 GB | Best open model, needs 64 GB+
-|===
-
-NOTE: `camel infra run ollama` runs Ollama in Docker and bypasses GPU 
acceleration,
-making inference significantly slower. Native install is preferred for 
development.
-
-NOTE: On Apple Silicon, use the default (GGUF) tags rather than the `-mlx` 
tags. The Ollama MLX
-engine cannot yet reuse the cached prompt for Qwen 3.x models, so every 
question re-processes the
-whole prompt, while the default engine reuses it and only processes what is 
new.
-
-=== Tool set for local models
-
-Every question sends the definitions of the `tui_*` tools the model may call, 
and a local model
-pays for each of them in prompt-processing time. The panel therefore sends 
only the core set of
-tools (state, tables, logs, errors, diagrams, topology, processor details, 
catalog docs, traces,
-spans, route control, sending messages, source files, infra services, 
navigation, log level and
-filters) to Ollama
-and to any provider on `localhost`, which roughly halves the prompt. Hosted 
providers get every
-tool, including the drawing, animation and automation tools. Use `/tools full` 
in the panel to send
-all tools to a local model too, `/tools core` to trim the set for a hosted 
one, pick *AI Tools* in
-*F2 -> Settings*, or set `camel.tui.ai.tools` in `.camel-cli.properties`. 
Ollama requests also ask the server to keep the
-model loaded for 30 minutes and for a context window of 32k or 64k (see 
<<_working_with_a_local_ollama_model>>;
-`OLLAMA_CONTEXT_LENGTH` overrides it), so follow-up questions reuse the cached 
prompt instead of reloading
-the model.
-
-=== Working with a local Ollama model
-
-A local model is not a slower version of a hosted one; it spends its time 
differently, and the TUI
-shows you where. This section explains what a question costs with Ollama and 
which knobs matter,
-with figures measured on an Apple M4 Pro (64 GB) running `qwen3.6:35b-a3b`.
-
-*How a question is spent.* Ollama answers in three phases, and every timing 
the TUI shows maps to one
-of them:
-
-1. *Load*: if the model is not in memory (first question, the keep-alive 
expired, or a request asked for
-   a different context size) Ollama starts a runner and loads the weights: 10 
to 20 seconds for a 22 GB
-   model. The TUI asks Ollama to keep the model loaded for 30 minutes after 
each request.
-2. *Prefill*: the prompt (system prompt, tool definitions, conversation 
history, your question) is
-   processed in one batch, at roughly 600 to 700 tokens per second when 
nothing is cached.
-3. *Decode*: the answer is generated token by token, at 50 to 60 tokens per 
second for this model.
-
-The wait before anything appears, the time to first token, is load plus 
prefill. A cold first question
-therefore takes 20 seconds before the first word; a warm one under a second.
-
-*One question, many requests.* The panel answers by calling the `tui_*` tools, 
and every tool call the
-model makes costs a new request that sends the whole prompt again. A simple 
question can take 3 to 13
-requests. This is affordable only because Ollama caches the prompt prefix: the 
system prompt, the tool
-definitions and the history are identical from one step to the next, so each 
step prefills only the new
-tokens and takes about half a second. Across a session the cache hit is 
typically above 90%. The
-Ollama tab shows the request count per question as `×N` and the cache hit per 
question; a question
-with a high count and a short answer is the model exploring, which a smaller, 
sharper tool set reduces
-(see <<_tool_set_for_local_models>>).
-
-*The context window.* The TUI decides the window it asks Ollama for 
(`num_ctx`) once per model:
-
-1. `OLLAMA_CONTEXT_LENGTH` in the environment wins when set.
-2. If the model is already loaded, its window is adopted (raised to 32k when 
smaller), so the TUI never
-   makes Ollama reload a model another client is using. A request with a 
different `num_ctx` costs a
-   cold start and throws away everyone's prompt cache.
-3. Otherwise 64k when the model's weights plus the KV cache of a 64k window 
fit the machine's memory,
-   computed from what `ollama show` reports (layers, KV heads, head size), 
else 32k. Overshooting is
-   worse than being conservative: Ollama then moves layers to the CPU and 
generation slows to a crawl.
-
-The static prefix of system prompt and core tools is about 4.5k tokens, and 
each question with tool
-calls adds another 2k to 4k of history. The panel compacts the history once 
the prompt Ollama measured
-for the last request passes half the window, capped at half of 64k even when a 
larger window was
-adopted: every token of history is prefill time again after a cache loss (an 
idle unload after 30
-minutes, another client, a restart), and 64k at a few hundred tokens per 
second is already well over a
-minute. Compacting rewrites the history, which invalidates Ollama's prompt 
cache, so the request after
-a compaction prefills the whole prompt again: about 40 seconds for a 21k-token 
prompt. The panel
-therefore compacts local history rarely and then thoroughly, down to about a 
quarter of the window,
-and prints one line saying what it did and how long the next reply will take 
to start. `/compact` does
-the same on demand with the same line; `/context` shows the window, the 
compaction point and the last
-measured prompt; the panel title shows the fill as `ctx 34%`.
-
-*Choosing the model.* With no `camel.tui.ai.model` set the panel takes 
`llama3.2` when it is installed,
-otherwise the first installed model; run `/model <name>` in the panel or set 
*AI Model* in
-*F2 -> Settings* to pin one. The model must support tool calling and should 
have at least 14B
-parameters; a mixture-of-experts model such as `qwen3.6:35b-a3b` prefills 
several times faster than a
-dense model of similar quality, which is what matters for a tool-heavy prompt. 
*F2 -> Run Doctor* shows
-whether Ollama was found, which models are installed and whether they are 
large enough.
-
-*Where to look.* The <<_ollama>> tab is the instrument for all of the above: 
tokens per second live and
-per request, time to first token with cold starts marked, cache hit, how full 
the context window is
-and how it grows per question, GPU and process load, and one line per question 
with the requests it
-took, the question being answered right now on top with `working` as its 
reason and its time counting
-up, and a footer with the average per question. In the AI panel, `/context` 
prints what the next request will cost, `/usage` and *Ctrl+U* the
-session totals per question.
-
-*Remote and containerised Ollama.* Everything above applies to an Ollama on 
another host or inside
-`camel infra run ollama` as well, with two differences: the container runs 
without GPU acceleration,
-and the live runner state and host load on the Ollama tab need the server on 
the same machine.
-
-For the wider picture see the blog posts
-link:/blog/2026/09/camel-local-model-benchmark/[We had a frontier AI coach a 
small local model through Camel]
-on what a local model can do with Camel and what was changed to help it, and
-link:/blog/2026/09/camel-genai-observability-jbang/[Observe Your Camel AI 
Routes with GenAI OpenTelemetry]
-on observing routes that call Ollama.
-
-=== Using an OpenAI-compatible local server
-
-Set `LLM_API_KEY` and `LLM_BASE_URL` to connect to any OpenAI-compatible server
-(LM Studio, vLLM, llama.cpp, GPT4All, …):
-
-[source,bash]
-----
-export LLM_API_KEY=any-value
-export LLM_BASE_URL=http://localhost:1234
-camel tui
-----
-
-`OPENAI_BASE_URL` is also accepted as an alternative to `LLM_BASE_URL`.
-
-The panel uses the first model the server lists on `/v1/models`. To use 
another one, run
-`/model <name>` in the panel (`/model` alone lists what the server offers), 
set *AI Model* in
-*F2 -> Settings*, or set `camel.tui.ai.model`. The model must support tool 
calling, otherwise the
-panel answers from training data instead of inspecting your integration. When 
a request fails, the
-panel shows the HTTP status and the server's error message, for example a 
model that the server
-does not host.
+To run the model on your own machine, with Ollama or an OpenAI-compatible 
server such as LM Studio, see
+xref:camel-jbang-tui-local-models.adoc[Camel TUI Local Models]: which models 
work well, the tool set they get,
+what a question costs and the Ollama tab.
 
 == AI panel slash commands
 
@@ -228,7 +76,7 @@ cycles backward.
 | Show what the next request costs: provider and model, tool set, static 
prefix size, history size and the session total, and with Ollama the context 
window, the prompt size above which the history is compacted and the last 
measured prompt. Useful with local models, where prompt size is time.
 
 | `/compact`
-| Shrink the conversation history sent to the model right away: older tool 
results are cut to their first lines and the oldest turns are dropped. With a 
hosted provider the panel does this automatically after each answer for all but 
the latest turn. With Ollama or another `localhost` provider it waits until the 
prompt Ollama measured passes half the context window (see 
<<_working_with_a_local_ollama_model>>), then compacts thoroughly, because a 
local server can reuse its cached prompt on [...]
+| Shrink the conversation history sent to the model right away: older tool 
results are cut to their first lines and the oldest turns are dropped. With a 
hosted provider the panel does this automatically after each answer for all but 
the latest turn. With Ollama or another `localhost` provider it waits until the 
prompt Ollama measured passes half the context window (see 
xref:camel-jbang-tui-local-models.adoc#_working_with_a_local_ollama_model[Working
 with a local Ollama model]), then comp [...]
 
 | `/retry`
 | Send the last question again, starting from a clean turn in the model 
history.
@@ -314,7 +162,7 @@ model called and the time spent in them, the number of 
requests, the tokens and,
 the context window got (for example `61.1s · 25 tool calls, limit reached 
(2.5s in tools) · 26 requests
 · 231.3k tokens · ctx 18%`). The panel allows 25 tool calls per question. When 
the model reaches that
 the byline says `limit reached`, the panel asks the model to answer from what 
it has found so far, and
-the <<_ollama>> tab shows the question's request count in red with `limit` as 
its reason (the count is
+the xref:camel-jbang-tui-local-models.adoc#_ollama_tab[Ollama tab] shows the 
question's request count in red with `limit` as its reason (the count is
 yellow from 10 requests, and the total time yellow from 30 seconds and orange 
from a minute). A question
 that runs into the limit is usually a tool that makes the model guess, not a 
weak model: the AI log
 shows which one.
@@ -353,45 +201,3 @@ on a 10k-token prompt means the cache was hit, while a 
prefill of several second
 prompt was processed again. For OpenAI, Anthropic and Gemini it prints the 
number of `cached` input
 tokens the provider reported. Use it to check that follow-up questions are 
cheap before blaming the
 model for being slow.
-
-== Ollama
-
-The Ollama tab (under More, in the *AI* group) shows how the model served by a 
local
-https://ollama.com[Ollama] is performing, in the spirit of an LLM dashboard. 
It works with or without a
-running integration: it finds Ollama at `localhost:11434`, at the address of 
`camel infra run ollama`,
-or at the endpoint the AI panel (*F8*) is using. The tab is listed only while 
an Ollama server answers;
-the TUI checks every ten seconds, so it appears shortly after `ollama serve` 
starts.
-
-* *Model* -- the loaded model with its family, parameters, quantization, 
layers, experts (and how many
-  are active per token for a mixture-of-experts model), how much of it sits in 
GPU memory, the
-  allocated context length and when Ollama will unload it. With no model 
loaded, the installed models
-  are listed instead.
-* *Throughput* -- decode and prefill tokens per second, *live* while the model 
is generating and
-  otherwise from the last request; time to first token and load time (a load 
of a second or more
-  is a cold start); session averages and a sparkline of the decode rate.
-* *Context* -- how full the context window is, the share of the prompt served 
from Ollama's cache,
-  whether the model is working or idle, the speculative decoding method in 
use, and a per-turn trend
-  of how much of the window each AI panel prompt filled, with the session peak 
and the compactions seen.
-* *Host* -- GPU utilization and memory (Apple silicon through `ioreg`, NVIDIA 
through `nvidia-smi`),
-  and CPU and memory of the Ollama server and its model runner.
-* *Requests* -- one line per question asked in the AI panel (a question with 
tool calls costs one
-  request per step; *Enter* unfolds the steps) with the question text, prompt 
and generated tokens,
-  cache hit, the share of the context window reached, prefill and decode 
tokens per second, time to
-  first token, the time you waited and the stop reason. Calls made by Camel 
routes are listed as
-  their own lines.
-
-Two kinds of requests appear in the log. Questions asked in the AI panel with 
Ollama as the provider
-come with the timings Ollama reports for each request (prompt evaluation, 
generation, model load,
-total). Calls made by Camel routes through `camel-langchain4j-chat`, 
`camel-openai` or
-`camel-spring-ai-chat` appear when the integration runs with GenAI 
observability (`--observe`, or
-`--dep=camel:ai-observability`), tagged with the route id; Camel records 
tokens and duration for those,
-not the phase split.
-
-The live figures, the context panel and the host panel need the model runner 
on the same machine:
-Ollama starts a `llama-server` process per loaded model and the tab reads its 
slot state a few times a
-second. Against a remote or containerised Ollama the tab keeps the model, 
per-request and session data
-and says which panels are unavailable.
-
-Press *r* to reset the request log and the session totals, *F5* to refresh 
immediately. The same data
-is available to AI agents through the `tui_get_ollama` MCP tool. For what the 
figures mean for the AI
-panel and which knobs to turn, see <<_working_with_a_local_ollama_model>>.
diff --git 
a/docs/user-manual/modules/ROOT/pages/camel-jbang-tui-getting-started.adoc 
b/docs/user-manual/modules/ROOT/pages/camel-jbang-tui-getting-started.adoc
new file mode 100644
index 000000000000..f7f854f0f5e9
--- /dev/null
+++ b/docs/user-manual/modules/ROOT/pages/camel-jbang-tui-getting-started.adoc
@@ -0,0 +1,124 @@
+= Camel TUI Getting Started
+
+There are several ways to get going with the TUI: point it at integrations 
that already run, run one of the
+built-in examples, open a project directory, or connect an existing Spring 
Boot or Quarkus application.
+
+See xref:camel-jbang-tui.adoc[Camel TUI] for the tabs, the keyboard shortcuts 
and the other pages.
+
+== Option 1: Your Own Route
+
+Start a Camel integration in one terminal:
+
+[source,bash]
+----
+camel run my-route.yaml
+----
+
+Open the TUI in another terminal:
+
+[source,bash]
+----
+camel tui
+----
+
+The TUI auto-discovers every running Camel integration on your machine -- no 
configuration needed.
+
+== Option 2: Built-in Examples
+
+Don't have a route yet? The TUI ships with a catalog of ready-to-run examples.
+Open the TUI and press *F2*, then select _Run an example_:
+
+[source,bash]
+----
+camel tui
+----
+
+. Press *F2* to open the actions menu
+. Select *Run an example...*
+. Browse the example catalog -- type to filter by name
+. Press *Enter* to launch the selected example
+
+Before launching, a run options form lets you choose the *runtime*: Camel Main 
(standalone),
+Spring Boot, Quarkus, or JBang. This makes it easy to try any example on all 
three runtimes without
+changing a single line of code. The first three run the example in a separate 
JVM that only
+contains the dependencies of the example (like a production deployment), while 
JBang runs it
+in-process in the Camel CLI JVM, which starts faster but has the CLI on the 
classpath as well. You can also set the integration name, toggle dev mode, and
+add extra dependencies. When the port the app will listen on (the one you set, 
else 8080) is taken already, the form
+says by whom, so a second app with an HTTP server does not fail to start.
+
+The example starts running in the background. The TUI auto-selects it as soon 
as it appears.
+From there you can explore tabs, watch messages flow, inspect the route 
diagram, and experiment.
+
+If an example requires infrastructure (like Kafka or a database), the TUI 
automatically starts the
+required Docker containers before launching the example. A notification in the 
footer shows the
+progress.
+
+The same runtime selector is available in *Run from folder...* (F2 menu), 
which lets you point
+the TUI at a local directory containing your routes. When a `pom.xml` is 
present, the runtime
+is auto-detected and locked to match your project.
+
+TIP: Press *F1* or *?* on any screen for context-sensitive help.
+Keyboard shortcuts are always shown in the footer bar.
+
+== Option 3: Open a Project Directory
+
+You can point the TUI directly at a project directory:
+
+[source,bash]
+----
+camel tui .
+camel tui /path/to/my-project
+----
+
+The TUI opens the directory in the Source tab so you can browse the project 
files immediately.
+When a `pom.xml` is present, the runtime is auto-detected (Spring Boot, 
Quarkus, or Camel Main).
+Press *F10* to run the project -- Maven projects are run with `camel run 
pom.xml`, which launches
+them with the appropriate goal (`spring-boot:run`, `quarkus:dev`, or 
`camel:run`), and plain
+directories are run with `camel run`. Either way the application logs to 
`~/.camel` so the Log tab
+shows its logs.
+
+Until it runs, the project is listed as *Stopped*: the Overview says what it 
is (runtime, Maven or route files, its
+folder) and the Source pane lists the routes it found and where they are, so 
*g* opens one. While it runs, the app
+takes its place under the name it gives itself, with the project's name next 
to it; when the run ends, the project is
+back as Stopped and still selected.
+
+This is a quick way to explore and run any Camel project without starting it 
separately first.
+
+image::jbang/camel-tui-main-quarkus-source.png[A Quarkus project in the Source 
tab: its Java route and the bean of the line]
+
+== Connecting Existing Spring Boot and Quarkus Applications
+
+The TUI auto-discovers integrations started with `camel run`. To monitor and 
control
+existing Spring Boot or Quarkus applications, add the `camel-cli-connector` 
dependency
+to your project. This is a lightweight runtime adapter that lets the TUI (and 
the Camel CLI)
+communicate with your application -- all TUI features work the same way 
regardless of runtime.
+
+Spring Boot:
+
+[source,xml]
+----
+<dependency>
+    <groupId>org.apache.camel.springboot</groupId>
+    <artifactId>camel-cli-connector-starter</artifactId>
+</dependency>
+----
+
+Quarkus:
+
+[source,xml]
+----
+<dependency>
+    <groupId>org.apache.camel.quarkus</groupId>
+    <artifactId>camel-quarkus-cli-connector</artifactId>
+</dependency>
+----
+
+Once added, start your application normally and the TUI will discover it 
automatically.
+No additional configuration is needed -- the connector auto-detects on the 
classpath and
+registers the application with the local Camel CLI.
+
+TIP: When you open a Spring Boot, Quarkus or Camel Main project via *F2 > Open 
Project* and run
+it with *F10*, the CLI connector dependency is automatically injected if it's 
not already in your
+`pom.xml`. This means the TUI can monitor the application without modifying 
your project.
+
+See xref:camel-jbang-managing.adoc[Managing Integrations] for more details.
diff --git 
a/docs/user-manual/modules/ROOT/pages/camel-jbang-tui-local-models.adoc 
b/docs/user-manual/modules/ROOT/pages/camel-jbang-tui-local-models.adoc
new file mode 100644
index 000000000000..aacdf2def9bf
--- /dev/null
+++ b/docs/user-manual/modules/ROOT/pages/camel-jbang-tui-local-models.adoc
@@ -0,0 +1,213 @@
+= Camel TUI Local Models
+
+The xref:camel-jbang-tui-ai.adoc[AI panel] (*F8*) works with a model running 
on your own machine, through
+https://ollama.com[Ollama] or any OpenAI-compatible server, so no API key is 
needed and nothing has to leave
+your machine. A local model spends its time differently from a hosted one, and 
the TUI shows you where.
+
+See xref:camel-jbang-tui.adoc[Camel TUI] for getting started and the other 
pages.
+
+== Key Features
+
+* <<_using_ollama_local_no_api_key,*Ollama*>> -- auto-detected at 
`localhost:11434`, with the models that work well
+* <<_using_an_openai_compatible_local_server,*OpenAI-compatible servers*>> -- 
LM Studio, vLLM, llama.cpp and others
+* <<_tool_set_for_local_models,*A smaller tool set*>> -- local models get the 
core tools, which roughly halves the prompt
+* <<_working_with_a_local_ollama_model,*What a question costs*>> -- load, 
prefill and decode, the prompt cache and the context window
+* <<_ollama_tab,*Ollama tab*>> -- the loaded model, tokens per second, cache 
hit and context fill, live
+
+== Using Ollama (local, no API key)
+
+Install Ollama natively for best performance — the native binary uses GPU 
acceleration
+(Metal on macOS, CUDA/ROCm on Linux):
+
+[source,bash]
+----
+# macOS
+brew install ollama
+
+# Linux
+curl -fsSL https://ollama.com/install.sh | sh
+
+# Pull a model — then open the TUI and press F8
+ollama pull qwen3.6:35b-a3b
+camel tui
+----
+
+Ollama at `localhost:11434` is auto-detected. No configuration needed.
+
+IMPORTANT: The F8 AI panel works by invoking built-in tools to inspect your 
running Camel process.
+Models smaller than ~14B do not reliably call tools and answer from training 
knowledge instead.
+Use at least a 14B model. Prefer a mixture-of-experts model such as 
`qwen3.6:35b-a3b`: with only
+3B parameters active per token it processes the tool-heavy prompt many times 
faster than a dense
+27B/32B model, so answers start in seconds instead of a minute.
+
+*Models that work well* (tool-calling capable, ≥14B, default Q4_K_M 
quantization):
+
+[options="header"]
+|===
+| Model | RAM | Notes
+| `qwen3.6:35b-a3b` | ~23 GB | Recommended: fastest prompt processing, needs 
32 GB+
+| `qwen2.5:14b` | ~9 GB | Minimum for 16 GB machines
+| `qwen3.6:27b` | ~18 GB | Strong dense model, several times slower prompt 
processing
+| `qwen2.5:32b` | ~20 GB | Good quality, slow prompt processing
+| `hermes3:70b` | ~43 GB | Excellent tool calling, needs 64 GB+
+| `llama3.3:70b` | ~43 GB | Best open model, needs 64 GB+
+|===
+
+NOTE: `camel infra run ollama` runs Ollama in Docker and bypasses GPU 
acceleration,
+making inference significantly slower. Native install is preferred for 
development.
+
+NOTE: On Apple Silicon, use the default (GGUF) tags rather than the `-mlx` 
tags. The Ollama MLX
+engine cannot yet reuse the cached prompt for Qwen 3.x models, so every 
question re-processes the
+whole prompt, while the default engine reuses it and only processes what is 
new.
+
+== Using an OpenAI-compatible local server
+
+Set `LLM_API_KEY` and `LLM_BASE_URL` to connect to any OpenAI-compatible server
+(LM Studio, vLLM, llama.cpp, GPT4All, …):
+
+[source,bash]
+----
+export LLM_API_KEY=any-value
+export LLM_BASE_URL=http://localhost:1234
+camel tui
+----
+
+`OPENAI_BASE_URL` is also accepted as an alternative to `LLM_BASE_URL`.
+
+The panel uses the first model the server lists on `/v1/models`. To use 
another one, run
+`/model <name>` in the panel (`/model` alone lists what the server offers), 
set *AI Model* in
+*F2 -> Settings*, or set `camel.tui.ai.model`. The model must support tool 
calling, otherwise the
+panel answers from training data instead of inspecting your integration. When 
a request fails, the
+panel shows the HTTP status and the server's error message, for example a 
model that the server
+does not host.
+
+== Tool set for local models
+
+Every question sends the definitions of the `tui_*` tools the model may call, 
and a local model
+pays for each of them in prompt-processing time. The panel therefore sends 
only the core set of
+tools (state, tables, logs, errors, diagrams, topology, processor details, 
catalog docs, traces,
+spans, route control, sending messages, source files, infra services, 
navigation, log level and
+filters) to Ollama
+and to any provider on `localhost`, which roughly halves the prompt. Hosted 
providers get every
+tool, including the drawing, animation and automation tools. Use `/tools full` 
in the panel to send
+all tools to a local model too, `/tools core` to trim the set for a hosted 
one, pick *AI Tools* in
+*F2 -> Settings*, or set `camel.tui.ai.tools` in `.camel-cli.properties`. 
Ollama requests also ask the server to keep the
+model loaded for 30 minutes and for a context window of 32k or 64k (see 
<<_working_with_a_local_ollama_model>>;
+`OLLAMA_CONTEXT_LENGTH` overrides it), so follow-up questions reuse the cached 
prompt instead of reloading
+the model.
+
+== Working with a local Ollama model
+
+A local model is not a slower version of a hosted one; it spends its time 
differently, and the TUI
+shows you where. This section explains what a question costs with Ollama and 
which knobs matter,
+with figures measured on an Apple M4 Pro (64 GB) running `qwen3.6:35b-a3b`.
+
+*How a question is spent.* Ollama answers in three phases, and every timing 
the TUI shows maps to one
+of them:
+
+1. *Load*: if the model is not in memory (first question, the keep-alive 
expired, or a request asked for
+   a different context size) Ollama starts a runner and loads the weights: 10 
to 20 seconds for a 22 GB
+   model. The TUI asks Ollama to keep the model loaded for 30 minutes after 
each request.
+2. *Prefill*: the prompt (system prompt, tool definitions, conversation 
history, your question) is
+   processed in one batch, at roughly 600 to 700 tokens per second when 
nothing is cached.
+3. *Decode*: the answer is generated token by token, at 50 to 60 tokens per 
second for this model.
+
+The wait before anything appears, the time to first token, is load plus 
prefill. A cold first question
+therefore takes 20 seconds before the first word; a warm one under a second.
+
+*One question, many requests.* The panel answers by calling the `tui_*` tools, 
and every tool call the
+model makes costs a new request that sends the whole prompt again. A simple 
question can take 3 to 13
+requests. This is affordable only because Ollama caches the prompt prefix: the 
system prompt, the tool
+definitions and the history are identical from one step to the next, so each 
step prefills only the new
+tokens and takes about half a second. Across a session the cache hit is 
typically above 90%. The
+Ollama tab shows the request count per question as `×N` and the cache hit per 
question; a question
+with a high count and a short answer is the model exploring, which a smaller, 
sharper tool set reduces
+(see <<_tool_set_for_local_models>>).
+
+*The context window.* The TUI decides the window it asks Ollama for 
(`num_ctx`) once per model:
+
+1. `OLLAMA_CONTEXT_LENGTH` in the environment wins when set.
+2. If the model is already loaded, its window is adopted (raised to 32k when 
smaller), so the TUI never
+   makes Ollama reload a model another client is using. A request with a 
different `num_ctx` costs a
+   cold start and throws away everyone's prompt cache.
+3. Otherwise 64k when the model's weights plus the KV cache of a 64k window 
fit the machine's memory,
+   computed from what `ollama show` reports (layers, KV heads, head size), 
else 32k. Overshooting is
+   worse than being conservative: Ollama then moves layers to the CPU and 
generation slows to a crawl.
+
+The static prefix of system prompt and core tools is about 4.5k tokens, and 
each question with tool
+calls adds another 2k to 4k of history. The panel compacts the history once 
the prompt Ollama measured
+for the last request passes half the window, capped at half of 64k even when a 
larger window was
+adopted: every token of history is prefill time again after a cache loss (an 
idle unload after 30
+minutes, another client, a restart), and 64k at a few hundred tokens per 
second is already well over a
+minute. Compacting rewrites the history, which invalidates Ollama's prompt 
cache, so the request after
+a compaction prefills the whole prompt again: about 40 seconds for a 21k-token 
prompt. The panel
+therefore compacts local history rarely and then thoroughly, down to about a 
quarter of the window,
+and prints one line saying what it did and how long the next reply will take 
to start. `/compact` does
+the same on demand with the same line; `/context` shows the window, the 
compaction point and the last
+measured prompt; the panel title shows the fill as `ctx 34%`.
+
+*Choosing the model.* With no `camel.tui.ai.model` set the panel takes 
`llama3.2` when it is installed,
+otherwise the first installed model; run `/model <name>` in the panel or set 
*AI Model* in
+*F2 -> Settings* to pin one. The model must support tool calling and should 
have at least 14B
+parameters; a mixture-of-experts model such as `qwen3.6:35b-a3b` prefills 
several times faster than a
+dense model of similar quality, which is what matters for a tool-heavy prompt. 
*F2 -> Run Doctor* shows
+whether Ollama was found, which models are installed and whether they are 
large enough.
+
+*Where to look.* The <<_ollama_tab,Ollama tab>> is the instrument for all of 
the above: tokens per second live and
+per request, time to first token with cold starts marked, cache hit, how full 
the context window is
+and how it grows per question, GPU and process load, and one line per question 
with the requests it
+took, the question being answered right now on top with `working` as its 
reason and its time counting
+up, and a footer with the average per question. In the AI panel, `/context` 
prints what the next request will cost, `/usage` and *Ctrl+U* the
+session totals per question.
+
+*Remote and containerised Ollama.* Everything above applies to an Ollama on 
another host or inside
+`camel infra run ollama` as well, with two differences: the container runs 
without GPU acceleration,
+and the live runner state and host load on the Ollama tab need the server on 
the same machine.
+
+For the wider picture see the blog posts
+link:/blog/2026/09/camel-local-model-benchmark/[We had a frontier AI coach a 
small local model through Camel]
+on what a local model can do with Camel and what was changed to help it, and
+link:/blog/2026/09/camel-genai-observability-jbang/[Observe Your Camel AI 
Routes with GenAI OpenTelemetry]
+on observing routes that call Ollama.
+
+== Ollama tab
+
+The Ollama tab (under More, in the *AI* group) shows how the model served by a 
local
+https://ollama.com[Ollama] is performing, in the spirit of an LLM dashboard. 
It works with or without a
+running integration: it finds Ollama at `localhost:11434`, at the address of 
`camel infra run ollama`,
+or at the endpoint the AI panel (*F8*) is using. The tab is listed only while 
an Ollama server answers;
+the TUI checks every ten seconds, so it appears shortly after `ollama serve` 
starts.
+
+* *Model* -- the loaded model with its family, parameters, quantization, 
layers, experts (and how many
+  are active per token for a mixture-of-experts model), how much of it sits in 
GPU memory, the
+  allocated context length and when Ollama will unload it. With no model 
loaded, the installed models
+  are listed instead.
+* *Throughput* -- decode and prefill tokens per second, *live* while the model 
is generating and
+  otherwise from the last request; time to first token and load time (a load 
of a second or more
+  is a cold start); session averages and a sparkline of the decode rate.
+* *Context* -- how full the context window is, the share of the prompt served 
from Ollama's cache,
+  whether the model is working or idle, the speculative decoding method in 
use, and a per-turn trend
+  of how much of the window each AI panel prompt filled, with the session peak 
and the compactions seen.
+* *Host* -- GPU utilization and memory (Apple silicon through `ioreg`, NVIDIA 
through `nvidia-smi`),
+  and CPU and memory of the Ollama server and its model runner.
+* *Requests* -- one line per question asked in the AI panel (a question with 
tool calls costs one
+  request per step; *Enter* unfolds the steps) with the question text, prompt 
and generated tokens,
+  cache hit, the share of the context window reached, prefill and decode 
tokens per second, time to
+  first token, the time you waited and the stop reason. Calls made by Camel 
routes are listed as
+  their own lines.
+
+Two kinds of requests appear in the log. Questions asked in the AI panel with 
Ollama as the provider
+come with the timings Ollama reports for each request (prompt evaluation, 
generation, model load,
+total). Calls made by Camel routes through `camel-langchain4j-chat`, 
`camel-openai` or
+`camel-spring-ai-chat` appear when the integration runs with GenAI 
observability (`--observe`, or
+`--dep=camel:ai-observability`), tagged with the route id; Camel records 
tokens and duration for those,
+not the phase split.
+
+The live figures, the context panel and the host panel need the model runner 
on the same machine:
+Ollama starts a `llama-server` process per loaded model and the tab reads its 
slot state a few times a
+second. Against a remote or containerised Ollama the tab keeps the model, 
per-request and session data
+and says which panels are unavailable.
+
+Press *r* to reset the request log and the session totals, *F5* to refresh 
immediately. The same data
+is available to AI agents through the `tui_get_ollama` MCP tool. For what the 
figures mean for the AI
+panel and which knobs to turn, see <<_working_with_a_local_ollama_model>>.
diff --git a/docs/user-manual/modules/ROOT/pages/camel-jbang-tui.adoc 
b/docs/user-manual/modules/ROOT/pages/camel-jbang-tui.adoc
index 2bb536b6a25d..78f90dddee3e 100644
--- a/docs/user-manual/modules/ROOT/pages/camel-jbang-tui.adoc
+++ b/docs/user-manual/modules/ROOT/pages/camel-jbang-tui.adoc
@@ -14,7 +14,7 @@ image::jbang/camel-tui-overview.png[TUI Overview showing 
multiple routes]
 == Key Features
 
 * *Run anything* -- your own routes, the built-in examples, or an existing 
Spring Boot or Quarkus project
-  (<<_getting_started,Getting started>>).
+  (xref:camel-jbang-tui-getting-started.adoc[Getting started]).
 * *Write routes with help* -- a xref:camel-jbang-tui-source-editor.adoc[source 
editor] that checks YAML, Java and XML
   routes as you type, completes endpoint options, fixes problems and shows the 
live run data of each line.
 * *See the integration* -- the xref:camel-jbang-tui-diagram.adoc[diagram] from 
the architecture down to one route,
@@ -28,126 +28,23 @@ image::jbang/camel-tui-overview.png[TUI Overview showing 
multiple routes]
 
 == Getting Started
 
-You can start using the TUI in two ways: with your own route, or by running 
one of the built-in examples.
-
-=== Option 1: Your Own Route
-
-Start a Camel integration in one terminal:
-
 [source,bash]
 ----
-camel run my-route.yaml
-----
-
-Open the TUI in another terminal:
-
-[source,bash]
-----
-camel tui
-----
-
-The TUI auto-discovers every running Camel integration on your machine -- no 
configuration needed.
-
-=== Option 2: Built-in Examples
+camel run my-route.yaml   # in one terminal
+camel tui                 # in another: auto-discovers every running 
integration
 
-Don't have a route yet? The TUI ships with a catalog of ready-to-run examples.
-Open the TUI and press *F2*, then select _Run an example_:
-
-[source,bash]
-----
-camel tui
+camel tui .               # or open a project directory and press F10 to run it
 ----
 
-. Press *F2* to open the actions menu
-. Select *Run an example...*
-. Browse the example catalog -- type to filter by name
-. Press *Enter* to launch the selected example
-
-Before launching, a run options form lets you choose the *runtime*: Camel Main 
(standalone),
-Spring Boot, Quarkus, or JBang. This makes it easy to try any example on all 
three runtimes without
-changing a single line of code. The first three run the example in a separate 
JVM that only
-contains the dependencies of the example (like a production deployment), while 
JBang runs it
-in-process in the Camel CLI JVM, which starts faster but has the CLI on the 
classpath as well. You can also set the integration name, toggle dev mode, and
-add extra dependencies. When the port the app will listen on (the one you set, 
else 8080) is taken already, the form
-says by whom, so a second app with an HTTP server does not fail to start.
-
-The example starts running in the background. The TUI auto-selects it as soon 
as it appears.
-From there you can explore tabs, watch messages flow, inspect the route 
diagram, and experiment.
-
-If an example requires infrastructure (like Kafka or a database), the TUI 
automatically starts the
-required Docker containers before launching the example. A notification in the 
footer shows the
-progress.
+No route yet? Press *F2* and pick _Run an example..._ to run one of the 
built-in examples on Camel Main,
+Spring Boot, Quarkus or JBang.
 
-The same runtime selector is available in *Run from folder...* (F2 menu), 
which lets you point
-the TUI at a local directory containing your routes. When a `pom.xml` is 
present, the runtime
-is auto-detected and locked to match your project.
+See xref:camel-jbang-tui-getting-started.adoc[Camel TUI Getting Started] for 
each way to start, and for
+connecting existing Spring Boot and Quarkus applications.
 
 TIP: Press *F1* or *?* on any screen for context-sensitive help.
 Keyboard shortcuts are always shown in the footer bar.
 
-=== Option 3: Open a Project Directory
-
-You can point the TUI directly at a project directory:
-
-[source,bash]
-----
-camel tui .
-camel tui /path/to/my-project
-----
-
-The TUI opens the directory in the Source tab so you can browse the project 
files immediately.
-When a `pom.xml` is present, the runtime is auto-detected (Spring Boot, 
Quarkus, or Camel Main).
-Press *F10* to run the project -- Maven projects are run with `camel run 
pom.xml`, which launches
-them with the appropriate goal (`spring-boot:run`, `quarkus:dev`, or 
`camel:run`), and plain
-directories are run with `camel run`. Either way the application logs to 
`~/.camel` so the Log tab
-shows its logs.
-
-Until it runs, the project is listed as *Stopped*: the Overview says what it 
is (runtime, Maven or route files, its
-folder) and the Source pane lists the routes it found and where they are, so 
*g* opens one. While it runs, the app
-takes its place under the name it gives itself, with the project's name next 
to it; when the run ends, the project is
-back as Stopped and still selected.
-
-This is a quick way to explore and run any Camel project without starting it 
separately first.
-
-image::jbang/camel-tui-main-quarkus-source.png[A Quarkus project in the Source 
tab: its Java route and the bean of the line]
-
-=== Connecting Existing Spring Boot and Quarkus Applications
-
-The TUI auto-discovers integrations started with `camel run`. To monitor and 
control
-existing Spring Boot or Quarkus applications, add the `camel-cli-connector` 
dependency
-to your project. This is a lightweight runtime adapter that lets the TUI (and 
the Camel CLI)
-communicate with your application -- all TUI features work the same way 
regardless of runtime.
-
-Spring Boot:
-
-[source,xml]
-----
-<dependency>
-    <groupId>org.apache.camel.springboot</groupId>
-    <artifactId>camel-cli-connector-starter</artifactId>
-</dependency>
-----
-
-Quarkus:
-
-[source,xml]
-----
-<dependency>
-    <groupId>org.apache.camel.quarkus</groupId>
-    <artifactId>camel-quarkus-cli-connector</artifactId>
-</dependency>
-----
-
-Once added, start your application normally and the TUI will discover it 
automatically.
-No additional configuration is needed -- the connector auto-detects on the 
classpath and
-registers the application with the local Camel CLI.
-
-TIP: When you open a Spring Boot, Quarkus or Camel Main project via *F2 > Open 
Project* and run
-it with *F10*, the CLI connector dependency is automatically injected if it's 
not already in your
-`pom.xml`. This means the TUI can monitor the application without modifying 
your project.
-
-See xref:camel-jbang-managing.adoc[Managing Integrations] for more details.
-
 == Tabs Overview
 
 The TUI organizes information into tabs. Press number keys *1* through *0* to 
jump directly
@@ -190,30 +87,22 @@ Two panels can be opened on top of any tab: *F6* opens an 
xref:camel-jbang-tui-s
 for running `camel` commands, and *F8* opens the 
xref:camel-jbang-tui-ai.adoc[AI panel] for asking
 questions about the running integrations.
 
-== Source Code Browser
-
-The Source tab (Tab 2) is a file explorer and editor for your project code 
that knows Camel. It checks your YAML,
-Java and XML routes as you type, completes endpoint options from the Camel 
catalog, explains the line the cursor is on,
-fixes problems for you (or asks the AI to), and shows the exchanges and 
failures of each line of the running
-integration next to the code.
-
-image::jbang/camel-tui-source-live-run-data.png[Live run data in the source 
editor]
-
-See xref:camel-jbang-tui-source-editor.adoc[Camel TUI Source Editor] for all 
its features and keys.
-
 == More about the TUI
 
 [cols="1,3",options="header"]
 |===
 | Page | What it covers
 
+| xref:camel-jbang-tui-getting-started.adoc[Getting Started] | Your own 
routes, the built-in examples, opening a project,
+connecting Spring Boot and Quarkus applications
 | xref:camel-jbang-tui-source-editor.adoc[Source Editor] | Reading and writing 
routes: checks as you type, quick fixes,
 fix with AI, completion, quick documentation, navigation, live run data
 | xref:camel-jbang-tui-diagram.adoc[Diagram] | The architecture, topology and 
route views, external endpoints, metrics
 | xref:camel-jbang-tui-observe.adoc[Observing Integrations] | Activity, 
message history, errors, spans, process, HTTP
 probe, CVE audit, Kafka, SQL, memory leaks, JFR, catalog
-| xref:camel-jbang-tui-ai.adoc[AI Panel] | AI providers and local models, 
slash commands, project overview, AI log, the
-Ollama tab
+| xref:camel-jbang-tui-ai.adoc[AI Panel] | AI providers, slash commands, 
project overview, AI log
+| xref:camel-jbang-tui-local-models.adoc[Local Models] | Ollama and 
OpenAI-compatible servers, the tool set for local
+models, what a question costs, the Ollama tab
 | xref:camel-jbang-tui-ai-agents.adoc[AI Agents] | MCP and ACP: what an AI 
agent can see and do, edits you confirm
 | xref:camel-jbang-tui-settings.adoc[Actions, Settings and Themes] | The 
actions menu, embedded shell, themes, settings,
 browser access, recording demos
@@ -276,7 +165,7 @@ See the 
xref:camel-jbang-tui-source-editor.adoc#_keyboard_shortcuts[keyboard sho
 | `100`
 
 | `--theme`
-| Color theme for this session (e.g., `dark`, `tokyo-night`, `dracula`). See 
<<Theme>> for the full list of 21 themes. Overrides the persisted 
`camel.tui.theme` preference when set.
+| Color theme for this session (e.g., `dark`, `tokyo-night`, `dracula`). See 
xref:camel-jbang-tui-settings.adoc#_theme[Theme] for the full list of 21 
themes. Overrides the persisted `camel.tui.theme` preference when set.
 |
 
 | `--record`

Reply via email to