Sorry, this email slipped by my inbox.

Martin Uecker <[email protected]> writes:

> Am Samstag, dem 09.05.2026 um 11:08 +0200 schrieb Arsen Arsenović:
>> Martin Uecker <[email protected]> writes:
>> 
>> > What I find unlucky, for example, is that we do have a generic
>> > tree data structure with a lot of imlicit assumptions and lot
>> > of issues if those are violated. So encapsulating this properly
>> > in abstract data types would be really helpful.  But we could do
>> > this also in C, while the switch to C++ apparently did nothing
>> > to address this.
>> 
>> Indeed, adding '-x c++' doesn't automatically fix existing code.
>> 
>> However, C++ enables other things to be done here.  For instance, at
>> one of the Office Hours, a mechanism for encoding tree types (without
>> requiring rewriting massive amounts of code) was proposed.  It relied on
>> templates, and it compiles down to the same code as today, and can
>> provide a bridge for existing code.  Essentially, it allowed us to
>> specify types such as 'ttree<one_of<PLUS_EXPR, MINUS_EXPR>>' or such,
>> and could automatically insert (and elide) gcc_asserts that check for
>> those tree codes.
>> 
>> I don't recall who proposed it or if it ever was sent, but a similar
>> thing wouldn't be possible in C.
>
> It isn't clear to me what that removing gcc_asserts by a complicated
> type (that then may be hidden behind "auto" because typing it out
> is to cumbersome) is necessarily an improvement.

It is an improvement.  The former detects errors at runtime, meaning
necessarily that it can only be performed if a codepath is reached (and,
what's more, not just if the assert itself is reached, but if the right
execution path is reached; not all paths that reach a given assert are
built the same of course).  The latter is a compile time error, meaning
incorrect paths cannot even be compiled (with the caveat that we'd
necessarily need to start with gradual typing, which, like typing.Any in
Python, would still permit *some* possibly incorrect code for
compatibility sake).

This exact distinction is the motivating drive between the "typed
qualifiers"[1] patch series I proposed recently (though, of course, that
one is a far smaller refactor that took less than a full workday to do
over the whole codebase after settling on a design, so it didn't need to
be done gradually like a change to 'tree' would).

Obviously, that case is far simpler, but that refactor - a miniature and
much simpler change compared to gradually-typed trees - managed to
detect dozens of bugs in the C++ FE while I was implementing support for
named address spaces.

It was well worth it, and I regret that I didn't do that before
embarking on implementing NAS support.  My intuition told me to do typed
qualifiers first (as soon as I started working on NAS support), but I
figured that I wouldn't have time to do it as part of the project that
required C++ NAS support.  In the end, delaying typed qualifiers cost
more time than implementing them.

I imagine the effects over the larger GCC codebase would be much
greater.  Though, because my (extremely unfinished) proposal is for the
introduction of effectually gradual typing, we probably won't see
extremely confronting immediate effects like with typed qualifiers, but
it should prevent a number of latent bugs and make code easier to
understand, and possibly detect bugs if someone was to spend some time
"propagating" typing through the codebase.

I'd love to have something to show to demonstrate what I was describing,
but I haven't had time to implement it yet.  It's not a very easy thing
to do all things considered.

> The issue with C++ is that people get overly excited about
> the impressive capabilities of the language (and don't get me wrong,
> there was a time in my life where I was too!), and in toy examples
> this always looks cool,  but then if you look at actual code in
> a larger project it is often much harder to understand.
>
> What we would need in my opinion would be simply
>
> struct expression;
> struct type;
> struct statement;
> etc.
>
> wrappers aroung tree nodes and a set of helpers function to operate on
> such data types without breaking their invariants, nothing more and
> nothing less.
>
> In C FE it looks somebody was starting to introduce such wrappers,
> but apparently never made a lot of progress. 

I wasn't proposing something complex in a vacuum or for fun.  I was
proposing something complex for a case where we have an essentially
untyped program, namely GCC, which is effectively dynamically typed
because 'tree' nodes are statically indistinguishable in the C and C++
type systems, where we should work to introduce static typing.

As experience with TypeScript and MyPy shows in the JavaScript and
Python ecosystems respectively, this requires a long time and, to be
feasible at all, requires the introduction of gradual typing constructs.

Gradual typing obviously doesn't exist in C and C++ OOTB.  However, it
can be implemented in the latter.

If we were writing GCC from scratch, having dedicated types instead of a
single 'tree' and some metaprogramming soup to implement a new type
system inside the C++ type system would easily be the better choice,
indeed, but we aren't writing GCC from scratch.

We have some three million lines of C++ code in GCC proper.  This
existing codebase frequently relies on patterns permitted by dynamic
typing, such as passing many different types as a single argument in a
single overload, sometimes even where they're of entirely different code
classes.

We do not have the luxury of picking something simple here.  We cannot
"just" do that refactor like I could with typed qualifiers.

Dynamic and gradual typing are the only choices.  The latter is easily
better.

On the "problem" of 'auto': the identifier 'tree' carries no semantic
meaning.  It can be anything.

Thus, it is strictly worse than 'auto', because 'auto' is a shorthand
for deduction.  It preserves the type of the value given to it, meaning
that the compiler, and tooling, like an LSP server for instance, and
thus editors, or even the user if a compiler diagnostic is emitted,
indeed do know the type.

'tree' erases the type of the value given to it.  Asserting on a 'tree'
verifies the type in one instant, but that verification is gone at time
of next def.

'auto' is the omission of redundant type information, but not its
erasure.

Of course, this redundancy can be useful because 1) the user truly does
specifically intend to constrain the type, where it matters that
e.g. they're only dealing with a specific tree_code(_class), or 2) a
reviewer looking at a patch does not have tooling immediately available
to them to reveal the type.

But, this can easily go the other way: if reason (1) isn't why the
redundancy is made explicit, then this just hinders the ability to
update some unrelated types in the codebase, because it requires usage
sites where this change isn't semantically significant to change.  A
declaration not using 'auto', that does not fall under (1), is
semantically over-constrained.

In my experience with codebases that rely heavily on type deduction
(most often in functional programming languages or in projects using the
aforementioned JavaScript and Python gradual typing extensions, or even
C++ codebases that IMO overuse 'auto'), the lack of the redundancy for
reason (2) is rarely a problem.

Thus, I advocate for using 'auto' in all cases except for those that
fall under (1).

[1] 
https://inbox.sourceware.org/gcc-patches/[email protected]/
-- 
Arsen Arsenović

Attachment: signature.asc
Description: PGP signature

Reply via email to