menu_open Columnists
We use cookies to provide some features and experiences in QOSHE

More information  .  Close

Training the Human Neural Network – RLHF to RLDF.

44 0
yesterday

How Deeds Trains the Human Neural Network

Before I begin this article, there is an interesting point here to note. Action always trumps knowledge. A recent post on the World Economic Forum on LinkedIn on marathon running echoes this and this article is partly inspired by that. I also ran the London Marathon 2014, so I sort of understand the runner’s mindset – but probably there are better ways to exercise – running is horrible for the knees especially on pavements (for me atleast). If you read this article and don’t act on it – then don’t expect to see the results from it. Praying can also be seen as quite relaxing, sort of like Yoga, for those into Yoga and Pilates, and this is some way to relate to it if you have never prayed before.

The second interesting point is that – this does not just apply to spirituality or religion, it applies to any domain of life – secular, moral, good or bad. Whatever you do, enjoy, feel rewarded by – will become habit. So make yours good habits – and feel rewarded by what you do of good and avoid the bad. Social media has “hijacked” the dopamine cycle in many ways – so we are being “trained” to spend hours doom scrolling on TikTok, YouTube Shorts or anything that gives “instant pleasure, gratification or novelty”. The “algorithm” for that training is “hidden” from you, and so are the negative effects on your brain, social life and moral (spiritual) life.

Thirdly, the video simply shows how these mechanisms work in the brain (as we understand them) and is only tangentially related to the article substance but to give readers an idea.

Artificial intelligence learns partly through feedback.

Give a system an objective. Let it act. Evaluate what it produces. Reinforce desirable behaviour. Penalise undesirable behaviour. Repeat the process often enough and behaviour changes.

In contemporary AI, RLHF normally means Reinforcement Learning from Human Feedback: humans express preferences about model outputs, and those preferences help shape subsequent behaviour.

Now reverse the picture.

What if the human being is the learning system?

Not literally a machine, of course. Islam gives humans consciousness, intention, moral responsibility and free choice. But as an analogy, reinforcement learning provides a surprisingly useful way of understanding the Islamic architecture of moral development.

Islam does not merely tell us what is true.

It creates a lifelong feedback system.

Good action is rewarded.

Bad action has consequences.

Errors can be corrected.

The complete training history is recorded.

And at the end comes an evaluation.

In machine-learning language, you might call it the ultimate human alignment problem.

The human neural network

A newborn human being does not arrive knowing how to conduct a business transaction, control anger, care for parents or give charity.

Human behaviour is progressively shaped.

Habits reinforce themselves.

The Qur’an and Sunnah add another layer: moral feedback linked to an ultimate objective.

Islam’s objective is not merely productivity or pleasure.

It is the formation of a human being who freely chooses what is good while recognising Allah.

The Qur’an describes this in the broadest possible terms:

“He who created death and life to test you as to which of you is best in deed.” — Qur’an 67:2

Life therefore contains something resembling a training environment.

You encounter wealth.

Opportunities to lie when nobody would know.

Opportunities to give when nobody would praise you.

The question is repeatedly the same:

The reward function is deliberately asymmetric

Here the Islamic model........

© The Times of Israel (Blogs)