The credit assignment problem, in the broader sense, refers to the
choice and validity of the exact attribution of "credit",
weight, factorization of  anything (and any *thing"):
feature, event, data point etc. as a cause, reason etc. for a
specific consequence. 


I guess you have not
read my comment or haven't parsed it properly or are ignoring my
reasoning.


My message argued
that the billionaires are only facades and avatars; they are not the
driving force, but only actors performing scripts, "puppets"
of the actual processes of the Universe which are deeper, broader,
distributed and reflected in many phenomena and events, not only in
these personas. "AI" is also a shorthand for many other
phenomena and events.


Yet your
interpretations are placing them in the center as drivers, actual
CCUs. 


AIXI/Hutter and also the referred 2012
paper by Bill, who cited Hutter, 2005 - just like Wolpert's theorems
and most mathematicians, Bostrom etc. - think in terms of *a
flat, single "utility function" - a "tyrant of
all", *without feedback, without interruptions, "master
of all".





IMO there is no such thing for partial
agents, subuniverses, virtual causality-control units.





"One" universal utility function may apply to the
Universe as a whole in some abstract sense and in some representation
of it, "The Master Algorithm" (Domingos) etc. and in some
selected abstract sense, view, interpretation at a selected
resolution of causality-control and perception, either as
prediction-compression, incremental predictive power or "maximum
pleasure" etc. However that is valid for every particular
causality-control unit, and there are multiple interpretations from
different views; also there are different definitions and
interpretations of "utilities" which can be seen as valid
at the same time. 









Intelligent systems with complex
behaviors are built from hierarchies, or are more effectively
represented and operated via hierarchies, of such CCUs, working at
different scales, ranges, domains , resolutions, precisions etc.
There is no  single "32-bit float value" that has to
be maximized. (There could be such, but only from a particular POV in
ML etc., however in order for the "loss function" to be
reduced, millions or billions of actual parameters have to be
modified and have to align in a distributed manner; the
evaluator-observer cherry picks the value of the "loss
function", while forgetting that the actual complexity lies in a
record that is billion times more complex just in raw bits, without
taking into account the inference and the relations between the bits;
see R.Ashby's *"Requisite Variety"; *poetically - *"The
loss function does not control the NN weights - the weights determine
the value of the loss function." *In addition: *The
structure of the system and its laws define the relations between the
weights and that computed value.*





There is no single "optimal"
behavior for complex multi-agent systems, as long as the system
continues to exist for a given range ahead or within some ranges of
possibilities, and a partial sub-universe cannot have a complete and
detailed model of the bigger or the whole Universe, as discussed in
regard to Wolpert's rediscoveries of these ideas presented in Theory
of Universe and Mind; therefore the smaller CCUs' predictions and
expectations are always "imperfect" and any of their
trajectories sooner or later will lead to a crash, i.e. the
optimality is always according to some "expectation", the
resolution is lower than the ultimate one - thus, variations are
allowed - and there is no absolute one.





Matt: *"happiness is the rate of
increase of utility."*





This is your definition; I disagree.
First, there is a limit: "bounded reward"; second: it
doesn't need to *increase* and the "happiness" for a simple
CCU - one that is not hierarchical, running at a single scale and
range; and happiness for  "sensual/physical", survival
CCUs - is not directly convertible to one which is constantly
"preempting", switching ranges, domains, virtual universes.
There is no single small CCU which "controls all" in
sophisticated systems, except when the system is malfunctioning or in
abstractions.





The "utility" and trajectory
of complex CCUs like humans, and the structures that are built
including the human individuals as "sub-agents" - the
society, the whole civilization, all humans, machines, technology,
data - are directed toward the development of more sophisticated
technology; curiosity and cognition. The individuals with a higher
intelligence and a proper neocortex-to-subcortical area projections
and "neural power" keep the lower-level utility function
under control and sometimes they completely override them; at least
sometimes or in some domains;  that applies at least in
individuals who are not slaves to the limbic system and the animal
and primitive type of CCUs, but their behavior and goals are
predominantly those involving cognitive value-free modules, not the
"lowest survival" layer of the dual system.





This topic was addressed even in the
beginning of Theory of Computers, Universe and Mind in the early
2000s, including the "endless pleasure", the goals of the
agents who "maximize utility".





In the lecture about the theory at the
world's first university course in AGI in 2010-2011, the difference
between these two types of CCUs is summarized in the following
slides:





https://research.twenkid.com/agi/2010/en/Todor_Arnaudov_Theory_of_Hierarchical_Universal_Simulators_of_universes_Eng_MTR_3.pdf





*See slides 18-23 (*they were included, but the platform seems to have problems 
with the images; hanging on Post etc. - if if "decides" to publish the previous 
attempt hours later - sorry*)

(...)

#20 

Reward/Success/Pleasure = ?*

● Indication of the degree of fulfiling the purpose of
existence/behavior of a given Control/Causality Unit


*Basic (Elementary) Purpose/Elementary Pleasure:*
Match between two values:

A) expected (desired/target) sensation (input)

B) actual sensation (input)


*Basic (elementary) input – number, variable:
*
IF (Input == Target_Input) Feels = NIRVANA;

IF (Input != Target_Input) Feels = HELL;

 *#21

 Graded Elementary Pleasure

*
● Distance, Difference, Comparison

● Animals and humans - pleasure/pain – indicate to
hypothetical CCU of a behavioral model how close the
subject is to the completion of the target state which is
initially preset: food, water, sex and of the anti-target:
hurting/pain/cold/hunger etc.


*Feels = Difference(Input, Target_Input)
*

Expected ~ predicted ~ desired (target)


The Difference between expected and reality is
a „mistake“ - displeasure.
Mind is aiming to decrease and eliminate the mistake

*#22*
*
*
*Cognitive and Physical pleasure*


*● Cognitive pleasure* – compression, prediction,
match, optimization (understanding, improving)


*Pleasure is Successful Prediction (of input).*


*Feels = Difference(Input, Predicted_Input)*


*
● Physical pleasure *– primary needs for survival in the
reality, the lowest (input) virtual universe. 
Primary
„rewards“ and purposes of behavior (self-preservation
(decrease pain), food, sex...)


*Pleasure is Desired Sensation (input).


Feels = Difference(Input, Desired_Input)*



*#23

Cognitive self-preservation „instinct“
is not really an instict*


● *Suicides* – believe (predict) they will feel less unpleasantly
after they kill themselves, than in the moment while they are
alive and making the decision to kill themselves, or after
executing some other behavior.

● Fear of death is not really [a] fear of death – see next slide

*● Physical and cognitive pleasure have to be
aligned/coordinated.*
● The cognitive hierarchy (Boris Kazachenko's term) or the
hierarchical simulators of virtual universes (Тodor Arnaudov)
receive feedback about the physical pleasure in the
environment where the mind exists; they are searching for
patterns/correlations between cognitive and physical
pleasure/reward.
*
*(...)


* I
think the camp of the "One-single-value-loss-function-rules-them
all" and other "flat-Earth"-ers wrongly  reduce
these two planes into one and interpret everything with a single number*; also 
these planes alone have many
subdimensions, resolutions, domains, precisions etc. as
mentioned.

 *The two reward-systems are connected and
there is an on-going process of their alignment, but the degree, the
details and the dynamics vary, depending on the exact structure,
distribution, weight and relations between the cognitive-sensual
CCUs, and there is preemption, interrupts, errors, gradual and abrupt
shifts, jumps etc.

** See the treatment of the deceptive "Liar's paradox" in *Universe and Mind 4, 
**it is about insufficient resolution of causality-control and perception; it 
is *reviewed also in *"***Wolpert's theorems about mutual 
unpredictability..."*, SIGI-2025 where D.Wolpert makes the same mistake of 
judging the Universe with flat binary "true/false" values.*
 
 Matt:* "*The
reason is that happiness is the rate of increase of utility.
Happiness is what you have minus what you want. We evolved to fear
death and then die, but what we really want is to stay in a stat**e** of 
maximum utility, which is the same thing. If a machine or drug
could give you the feeling of a million orgasms for as long as you
wanted, what do you think you would do?"
 
 *Even
for the sensual CCUs ("physical" in the slides?), it is
enough if they "stay", as you also say.
 
 Humans
have different pleasures and not everybody is a slave to the one that
you mention. Yes, the most basic one and the one which is "sold"
to the masses by the religions is namely that: "eternal
pleasure" vs "eternal displeasure", Paradise and Hell
and the "orgasms" are again suggested as some "ideal"
pleasure. Perhaps drug addicts may propose "having a shot of..."
their preferred narcotic.
 
 However in the real world for
humans of flesh and blood, who are not "hijacked" by such a
bug, *this
actually doesn't work.
 
 *Behaviors
which manifest extreme addictions are an error for the complex and
especially *cognitively
dominant* CCUs,
or at least ones which sometimes have cognitive dominance - these
CCUs are being locked in an endless cycle by some coalition of
sub-units, and that cycle is too short and primitive; the fault is
evident to similar-class external evaluator-observers, who realize
this is inappropriate- drug addictions, obsessions,
"obsessive-compulsive disorder" etc. 
 
 Animals
prefer food to this kind of pleasure, sexual reproduction is a
"job".
 
 The human sexual drive is also likely a
biological short-circuit that exists in order to force humans to
reproduce; if the cognitive system is free, it may prefer cognitive
or more complex "pleasures"  too often - as it happens
with many intellectuals*.  Evolutionarily, the origin of the
neocortex is from a brain area dedicated to controlling the sexual
behavior in reptiles, according to a remark by S.Savelyev. That
origin may explain some "bugs" in the human brain, the
over-drive on sexual activities, while having much higher "requisite
variety" for other domains.
 
 Yet, humans are not so
much sexually driven as to want to have orgasms all day and all
night. 



The
pop-culture often exploits and ridicules the plot of a man being
unable to "ride the horse" for long enough and the
"unsatisfied women" - see for example the sitcom "Married
with Children".
 
 However the truth is that "normal"
women do not like men who can last for too long either, so long as
they have already reached orgasm, unless they are nymphomaniacs. 
 
 If
it can be done in 20 or 30 minutes - that is enough. If you run for
an hour or 1:30 hours to finish - "healthy" women get bored
or tired and they do not take it as "more pleasure". It is
also a stress on the body, the tissues on both sides are
tender*.
 
 Humans have many other kinds of pleasure and
drives and utility functions,  even at the limbic and animal
level, and they have one general "displeasure" function: *boredom*.
"Orgasm" is only one dimension.
 
 For example food
preference is illustrated in an episode of the sitcom *"Everybody
loves Raymond"* - some of the late seasons; the old father Frank secretly 
starts to
use pills for his masculinity, in order to preserve his sexual life;
however a side effect of the pill is a loss of sense of taste and
smell, which makes him stop enjoying the food cooked by his wife
Mary; that creates tension between them; in the end, after the
problem is unveiled, Frank  decides that he prefers to enjoy the
food that his wife prepares and both are happy.
 
 As a
common story goes, some kind of fighters are attracted to sacrifice
their lives, because they "will meet 100 virgins in Paradise"
[change the number and the kind of women with your desires: "barely
legal teens, MILFs, big boobs, secretaries, ..."]. 
 
 If
this is a true motivation at all, and not propaganda nonsense, the
ones who "buy" this perhaps had some mental or "biological"
issues from the start, and possibly there were *social
issues* 
- e.g. they cannot find a wife, due to wars, destroyed social
structure and relations etc., and this is how the men are
recruited.
  
 The average believers in stable societies
from the same major religion do not "take the bait" after
such invitations; men with a wife and children do not go for the
virgins and leave their families. The believers in the other major
religion had similar  "calls for action" and also went
for Crusades, but again the ones who take it were a specific type and
the pretended goal was likely a nonsense as well, similar to modern
declarations of wars for "democracy, human rights" etc.
while everybody knows the real goals are just power, resources, land,
domination, "to win".
 
 Another
low-level example is dopamine, with which "everything" is
explained. Food and sugar in particular raise dopamine, "sugar
acts like cocaine". It is pleasant, "you want to eat more",
return to eating ice cream, chocolate bars etc.




Is it really so, though? In the context
of your definition: 





*Do healthy humans want more and more
sweet foods? "Increasing the utility function"?*





- Not really.





There is feedback and negative
feedback, "boredom" and variety of foods - at least in the
society where there is abundance of food.





If the organism is intact, you may eat
a croissant, a waffle, some chocolate, some ice cream, but you will
get satisfied, and then eating more will be unpleasant, you're
already full. If you eat too often a particular candy, it begins to
be "boring", "let's try something else". In
the long run there would be other negative-feedback signals. 





Sometimes, if you need more, it is
because the glycogen stores are depleted and the organism, the
virtual complete CCU of the body, requires that these reserves are
replenished as soon as possible.





If the person really needs more and
more sugar, in order to get satisfied, or he never gets satisfied
("endless orgasm"*), that is a sign of malfunction -
insulin resistance, prediabetes, diabetes.





....





* The researchers from the school of *"Relevance Realization" *understand that 
there
is a hierarchy of nested CCUs and they call this "incomputable",
because they base their notion of computation on the Turing machine,
which is also flat. However the "Universe Computer" and the
computers do not have to be like that and in fact the real computers
do not operate as a simple Turing Machine, they "can be
represented by one". See a discussion and review of the
literature in the appendix *"Algorithmic Complexity"* in *The Prophets. *





See a review in "The Prophets...",
volume "Listove: Reflections on Everything" and a specific
paper, comparisons of passages, made by  LLMs at SIGI-2025.


https://github.com/Twenkid/SIGI-2025/blob/main/LLM/Relevance-Realization-21-10-2025-ChatGPT5.pdf



https://github.com/Twenkid/SIGI-2025/blob/main/LLM/Relevance-Realization-TOUM-Correspondences-21-10-2025-Kimi_2-Claude4_5-ChatGPT5.pdf







* See also the research in
neuroscience about the *"Winnerless competition" *of
the networks in the brain, by M.Rabinovich et al. ; Karl Friston's
work is related as well; there is a review of the literature and
discussion in the mentioned volume "Listove".





* See also the criticism on the
"Marshmallow test" on Artificial Mind and the extended
version of the article in the main volume of *The Prophets: *


https://artificial-mind.blogspot.com/2018/06/delayed-gratification-is-ill-defined.html


*In Analysis, Articles, Developmental
Psychology, Neuroscience by Todor "Tosh" Arnaudov - Twenkid
// Wednesday, June 27, 2018 // Leave a Comment*


** Delayed gratification is an
ILL-Defined Concept*


In the context of this conversation, *"two marshmallows", *either after the 
adult
returns or at the moment, are not strictly "a larger
reward"  or provide "more pleasure" for the child
and they do not represent by default a bigger value of the "utility
function". This is only *suggested *by the experimenters.
The test is of obedience to authority, "patience" etc. The
child may have many other "rewards", may not see two items
as "more" (the taste is the same and both may be eaten in a
short window of time), she may prefer other types of food, may be
satisfied with one and immediately go and do and experience something
else: play, eat another food, meet her mother etc.





Recently it was  discovered that
the parents of the children who do not wait for the second
marshmallow, often break their promises; therefore the children may
not trust the experimenters, even if they wanted a second candy.


* Regarding motivation of
intellectuals, see
the idea satirized in the movie "Idiocracy", 2006.





------------------------------------------
Artificial General Intelligence List: AGI
Permalink: 
https://agi.topicbox.com/groups/agi/Tc0ac46f2c46f6874-Ma331510f96d1e3d2e2d3312f
Delivery options: https://agi.topicbox.com/groups/agi/subscription

Reply via email to