Matt Turner <[email protected]> writes:

> The per-CPU TB jump cache has held 4096 entries since it was introduced.
> That is too small for guests running large programs: an emulated compiler
> misses often enough that the fallback qht lookup shows up prominently in
> a profile.
>
> Measured with qemu-alpha running an emulated alpha gcc 16.2.0 compiling
> the SQLite 3.45.1 amalgamation (255k lines, -O2) on an x86-64 host. The
> compile performs 34.2 billion TB executions, of which 8.4 billion take
> the indirect dispatch path.
>
> Sizing curve, on top of the preceding patch, instructions retired and
> wall clock:
>
>     12 bits (  64 KiB): 1,563,829,403,943          133.13s
>     14 bits ( 256 KiB): 1,493,865,985,972  -4.47%  124.89s  -6.19%
>     16 bits (   1 MiB): 1,469,729,281,442  -6.02%  120.97s  -9.13%
>     18 bits (   4 MiB): 1,462,262,363,257  -6.49%  120.16s  -9.74%

I'd be curious to see the system emulation numbers. You should see very
different profiles as more stuff gets directly chained in user mode.

> 16 bits is the knee. 18 buys another 0.47% of instructions for four times
> the memory, and since instructions retired does not account for the data
> cache pressure of a 4 MiB table, that 0.47% is probably not real: the
> wall clock difference between 16 and 18 bits is 0.67%, against a
> run-to-run spread of the same order.
>
> In a perf profile the mechanism is visible directly: tb_htable_lookup(),
> which is where qht_lookup_custom() lands once it is inlined in an LTO
> build, falls from 5.73% of samples to 1.52%.
>
> The cost is memory: the cache grows from 64 KiB to 1 MiB per vCPU. That
> is easy to justify for a single-vCPU linux-user process and less obvious
> for system emulation with many vCPUs, so this may want to be sized by
> target or made tunable rather than raised unconditionally.

Practically you wouldn't see many TCG emulations with more the 16 CPUs
(although it would be interesting to see where the MTTCG gains fade with
these proposals). With that in mind 1MiB doesn't see too excessive - I
don't know how it compares to other big allocations QEMU makes over a
typical run.

>
> Signed-off-by: Matt Turner <[email protected]>
> ---
>  accel/tcg/tb-jmp-cache.h | 2 +-
>  1 file changed, 1 insertion(+), 1 deletion(-)
>
> diff --git ./accel/tcg/tb-jmp-cache.h ./accel/tcg/tb-jmp-cache.h
> index c3a505e394..268dacd7ba 100644
> --- ./accel/tcg/tb-jmp-cache.h
> +++ ./accel/tcg/tb-jmp-cache.h
> @@ -12,7 +12,7 @@
>  #include "qemu/rcu.h"
>  #include "exec/cpu-common.h"
>  
> -#define TB_JMP_CACHE_BITS 12
> +#define TB_JMP_CACHE_BITS 16
>  #define TB_JMP_CACHE_SIZE (1 << TB_JMP_CACHE_BITS)
>  
>  /*

-- 
Alex Bennée
Virtualisation Tech Lead @ Linaro

Reply via email to