On Mon, Jun 22, 2026 at 9:32 PM James Bowery <[email protected]> wrote:
> On Mon, Jun 22, 2026 at 7:21 PM Matt Mahoney <[email protected]> wrote:
>> I agree. We need a definition of friendly (or aligned) AI that we can all 
>> agree on. What is your definition?
> You missed the point of my using the term "Leggian".

Fair enough. Let me make a list of "friendly AI" definitions and find
an elegant mathematical model that's consistent with most of them.

1. Asimov's 3 laws of robotics (1942).
https://en.wikipedia.org/wiki/Three_Laws_of_Robotics

2. Yudkowsky's Coherent Extrapolated Volition (2004).
https://intelligence.org/files/CEV.pdf

3. Various modern LLM directives intended to avoid government
regulation and/or lawsuits. Don't be racist. Don't generate nude
images of celebrities. Don't encourage users to commit suicide. Don't
tell users how to create biological weapons.

4a/b. Effective Altruism: Maximize collective utility. If life has
positive utility (because death is bad), then life should proliferate
to the maximum extent possible. If life has negative utility (because
all living things must die), then all life should be exterminated now
to eliminate future suffering.

A common theme is more happiness and less death and suffering among
users. But this won't work. I pointed out earlier that happiness is
the rate of increase of utility. But we live in a finite universe that
can only support at most 10^92 bit write operations before the heat
death of the universe and everything in it. Therefore all utility
functions are finite and have one or more maximum states. The best
that a goal seeking agent can do is reach one of these states and stay
there. That is equivalent to death.

Furthermore, suffering is not the same as pain or negative
reinforcement. Pain is a signal that alters your memory to make you
fear the sensory perceptions associated with it. You interpret this
fear as a memory of suffering to be consistent with your illusion of
free will.

Intuitively, a reinforcement signal results in some change in
behavior, and stronger signals (more intense pain or pleasure) result
in greater changes. We can give an algorithmic definition of the
strength of a reinforcement signal. Let S1 be the state of the agent
before the signal is applied, and S2 the state after. Then the
strength of the reinforcement signal is the conditional Kolmogorov
complexity K(S2|S1), or the length of the shortest program that
describes S2 given S1 as input.

For example, my simple 2007 reinforcement learner Autobliss (
https://mattmahoney.net/autobliss.txt ) is a programmable 2 input
logic gate programmed by reinforcement learning. It has a 4 bit state,
so the most pain or pleasure it can experience between states is 4
bits. It could experience more by being programmed to different
configurations, switching every few nanoseconds, but this erases the
memory of earlier experiences. Human brains have a conscious memory
capacity of 10^9 bits and a write speed of 5 to 10 bits per second, so
our capacity to experience pain and pleasure has a much lower rate,
while our total experience is greater.

An obvious drawback of my definition is that it does not distinguish
between positive and negative reinforcement. That's because in a pure
reinforcement learner, it doesn't matter. But in biological organisms,
positive reinforcement signals generally contribute to reproductive
success, while negative signals are detrimental. Furthermore, social
animals including humans signal their mental states in a way that
negative signals cause distress among other organisms. We call this
phenomenon "empathy". Without empathy, human civilization would not
have happened. I model this in Autobliss by having it signal "Ahhh" or
"Ouch" and killing it if it receives too much negative reinforcement.

We are almost ready to give a mathematical description of friendly AI.
Without reference to utility, we want more positive reinforcement than
negative reinforcement. This implies the proliferation of life
provided we exclude destructive positive reinforcement such as by
drugs or wireheading. This is consistent with definitions 1, 2, 3, and
4a (but not 4b). The first law of robotics says not to harm humans.
CEV gives you what you would want if you were smarter, not everything
you want right now. We want everyone to have this because empathy
leads to cooperation.

Note also that K(death) = K(death|S1) = 0. Death does not cause pain
or suffering. Therefore we may reject 4b.

The definition of life is self replication. It doesn't have to be by
binary fission or sexual reproduction. It could involve
specialization, like a queen ant or robots building factories that
build robots. The key step is copying information, which decreases the
entropy of the atoms that make up the copy, which requires increasing
the entropy somewhere else, which requires free energy, of which 10^70
J or 10^92 bits are available in the observable universe.

Therefore, I define the friendliness of AI to be the rate of entropy
removed by replicating entities, measured in bits per second. A bit is
at least kT ln 2, or about 3 x 10^-21 J at room temperature.

I'm not sure how useful this is. It measures how much life there is,
with the assumption that more is better. It is weighted toward active
life, animals over plants. But it also measures industrial activity,
weighted toward energy efficiency. Manufacturing moves atoms, which
reduces entropy and takes energy just like biological processes.
Making steel to build robots means separating iron atoms from oxygen
atoms.

The Earth intercepts 177,000 TW of sunlight, of which 90,000 TW
reaches the ground. Plants make 300 TW of food from sunlight, of which
1 TW is consumed by humans. In addition we produce 20 TW of fuel
energy, of which 3.5 TW is converted to electricity. These numbers are
growing, so we are headed in the right direction.

------------------------------------------
Artificial General Intelligence List: AGI
Permalink: 
https://agi.topicbox.com/groups/agi/Teedaf4a19e56877e-Mad8c7c5829ea3a13d1b411a4
Delivery options: https://agi.topicbox.com/groups/agi/subscription

Reply via email to