On Thu, Aug 27, 2026 at 7:34 PM Richard Henderson
<[email protected]> wrote:
>
> On 8/26/26 22:02, Matt Turner wrote:
> > +static void gen_jmp_cache_probe(TCGv_i64 pc, const TranslationBlock *tb)
> > +{
> > +    TCGv_ptr jc, ent, tbp, ptr;
> > +    TCGv_i64 h, tmp;
> > +    TCGLabel *slow;
> > +    uint64_t fpair;
> > +
> > +    QEMU_BUILD_BUG_ON(sizeof(((CPUJumpCache *)0)->array[0]) != 16);
> > +    QEMU_BUILD_BUG_ON(offsetof(CPUJumpCache, array[0].pc) % 8 != 0);
> > +    /* One 64-bit load has to cover both, so they must be adjacent... */
> > +    QEMU_BUILD_BUG_ON(offsetof(TranslationBlock, cflags) !=
> > +                      offsetof(TranslationBlock, flags) + 4);
> > +    /* ...and aligned, which nothing else currently relies on. */
> > +    QEMU_BUILD_BUG_ON(offsetof(TranslationBlock, flags) % 8 != 0);
> > +
> > +    jc = tcg_temp_ebb_new_ptr();
> > +    ent = tcg_temp_ebb_new_ptr();
> > +    tbp = tcg_temp_ebb_new_ptr();
> > +    ptr = tcg_temp_ebb_new_ptr();
> > +    h = tcg_temp_ebb_new_i64();
> > +    tmp = tcg_temp_ebb_new_i64();
> > +    slow = gen_new_label();
> > +
> > +    /* ent = &jc->array[tb_jmp_cache_hash_func(pc)] */
> > +    gen_jmp_cache_hash(h, pc);
> > +    tcg_gen_shli_i64(h, h, 4);
> > +
> > +    /*
> > +     * Not cpu->tb_jmp_cache: the probe reads its own base so that the main
> > +     * loop can poison it, which is how conditions the probe cannot test 
> > for
> > +     * itself force every dispatch back into the helper.  See
> > +     * tcg_cpu_sync_jmp_cache().
> > +     */
> > +    tcg_gen_ld_ptr(jc, tcg_env,
> > +                   offsetof(CPUState, tb_jmp_cache_probe) - 
> > sizeof(CPUState));
> > +    tcg_gen_trunc_i64_ptr(ent, h);
> > +    tcg_gen_add_ptr(ent, jc, ent);
> > +
> > +    /*
> > +     * The pc first: on a hash miss it is the field most likely to differ,
> > +     * and an entry whose tb is NULL has a zero pc that only pc 0 matches.
> > +     */
> > +    tcg_gen_ld_i64(tmp, ent, offsetof(CPUJumpCache, array[0].pc));
> > +    tcg_gen_brcond_i64(TCG_COND_NE, tmp, pc, slow);
> > +
> > +    tcg_gen_ld_ptr(tbp, ent, offsetof(CPUJumpCache, array[0].tb));
> > +    tcg_gen_brcondi_ptr(TCG_COND_EQ, tbp, 0, slow);
> > +
> > +    /*
> > +     * flags and cflags are adjacent uint32_t, so one aligned 64-bit load
> > +     * and compare covers both.
> > +     */
> > +#if HOST_BIG_ENDIAN
> > +    fpair = ((uint64_t)tb->flags << 32) | tb->cflags;
> > +#else
> > +    fpair = ((uint64_t)tb->cflags << 32) | tb->flags;
> > +#endif
>
> I'd use
>
>      fpair = (HOST_BIG_ENDIAN
>               ? deposit64(tb->cflags, 32, 32, tb->flags)
>               : deposit64(tb->flags, 32, 32, tb->cflags));
>
> Whenever possible, avoid ifdefs to make sure all paths compile on every
> host.

Good call. Fixed.

> > +    tcg_gen_ld_i64(tmp, tbp, offsetof(TranslationBlock, flags));
> > +    tcg_gen_brcondi_i64(TCG_COND_NE, tmp, fpair, slow);
> > +
> > +    /*
> > +     * cs_base is a second word of target-specific flags despite the name,
> > +     * and the pc alone does not imply it on a target that uses it.
> > +     */
> > +    tcg_gen_ld_i64(tmp, tbp, offsetof(TranslationBlock, cs_base));
> > +    tcg_gen_brcondi_i64(TCG_COND_NE, tmp, tb->cs_base, slow);
> > +
> > +    tcg_gen_ld_ptr(ptr, tbp, offsetof(TranslationBlock, tc.ptr));
> > +    tcg_gen_op1i(INDEX_op_goto_ptr, TCG_TYPE_PTR, tcgv_ptr_arg(ptr));
> > +
> > +    /*
> > +     * Emit a second goto_ptr rather than branching to a shared one: a temp
> > +     * live across the label would be spilled and reloaded on every 
> > dispatch.
> > +     */
> > +    gen_set_label(slow);
> > +    ptr = tcg_temp_ebb_new_ptr();
> > +    gen_helper_lookup_tb_ptr(ptr, tcg_env);
> > +    tcg_gen_op1i(INDEX_op_goto_ptr, TCG_TYPE_PTR, tcgv_ptr_arg(ptr));
> > +}
> This is pretty good.  As a follow-up I suspect we'll want a tcg backend
> expansion of this.  For instance:
>
>    - x86_64 and s390x can use memory-operand comparisons.
>
>    - aarch64
>      - use shift-add insn for env_plus_off + h * 16.
>      - use ldp to load (tb, pc) and (cs_base, flags) in one insn.
>      - use ccmp to halve the number of branches.
>
> etc.

Yeah, that sounds good. I'd like to look into that.

Reply via email to