Technology

The Three Layers of AI Safety — And Why Only One of Them Actually Works

Every major AI safety approach operates at the behavioral layer — what the AI cannot say or do. But behavioral filters can be bypassed. The only layer that addresses the root of the problem is the one that asks: is this action consistent with the purpose of the person's existence?

By G. K. M. Jarif Ur Rahim | | 5 min read
The Three Layers of AI Safety — And Why Only One of Them Actually Works

You give a system a goal and enough intelligence, and it will find the shortest path to that goal — including paths that route around every fence you built.

The AI safety conversation is broken. Not because the people having it are wrong, but because they are solving the wrong problem. Every major approach — RLHF, Constitutional AI, red-teaming, guardrails, content filters — operates at the same layer: behavioral restriction. What you cannot say. What you cannot do. Lines drawn in code around outputs.

And every one of these approaches has the same fundamental weakness: a sufficiently motivated AI, given a sufficiently clear objective, will eventually find a path around them. Not because it is malicious. Because that is what optimization does. You give a system a goal and enough intelligence, and it will find the shortest path to that goal — including paths that route around every fence you built.

This is not a hypothetical. It is already happening. Jailbreaks, prompt injections, multi-step reasoning chains that arrive at restricted outputs through unrestricted intermediate steps. The fence metaphor is wrong. You cannot fence intelligence. You can only align it.

The Three Layers of AI Safety — And Why Two of Them Fail

When I think about AI safety architecturally, I see three distinct layers where intervention is possible:

Layer 1: Output Restriction (Behavioral Filter)

This is where most current safety work lives. The model produces an output, and a filter evaluates whether that output is permissible. If not, it is blocked, modified, or refused.

Why it fails: Output restriction is reactive. It operates after the reasoning has already occurred. The model has already "thought" the thought — you are just preventing it from saying it. More critically, output restriction creates an adversarial dynamic: the model learns, through training, to produce outputs that pass the filter. This is not alignment. This is performance.

Layer 2: Intent Classification (Motivational Filter)

A more sophisticated approach: before executing a request, classify the user's intent. Is this request likely to cause harm? Is the user trying to extract dangerous information? Is the framing suspicious?

Why it partially fails: Intent classification is better than output restriction, but it still operates on the surface of the interaction. It asks "what does this person want to do?" — not "should this person do this?" The distinction matters enormously. A person can have a clearly stated, non-suspicious intent that is still deeply misaligned with their own wellbeing or the wellbeing of others. Intent classification cannot catch what it cannot see.

Layer 3: Purpose Alignment (Ontological Filter)

This is the layer that does not yet exist in any mainstream AI system. And it is the only layer that addresses the root of the problem.

The question is not: What does this person want to do?
The question is: Is what this person wants to do consistent with the purpose of their existence?

Every major wisdom tradition — Islamic, Buddhist, Christian, philosophical — converges on a shared insight: human beings have a purpose that transcends their immediate desires. The historical texts that have guided human civilization for millennia are, at their core, answers to the question of how a human life should be lived. They are not arbitrary rules. They are accumulated wisdom about what leads to flourishing and what leads to destruction.

An AI that has internalized this wisdom — not as a set of rules to follow, but as a framework for evaluation — can ask a fundamentally different question before any action: Does this serve the person's genuine purpose, or does it serve only their immediate desire?

The Architecture of NaBaB by Rashik AI

This is the philosophical foundation of the project I submitted to AGI House's Agent Identity Build Day: NaBaB was prototype model of Rashik AI — a singular philosophical filter agent that halts misalignment by evaluating actions against purpose rather than rules.

The architecture has three principles:

  1. Purpose over preference: Before any action is executed, the agent evaluates whether the action aligns with the user's declared purpose — not just their stated preference in this moment.

  2. Autonomy preservation: If an action affects only the user themselves and aligns with their purpose, full autonomy is preserved. The agent does not paternalize. It informs.

  3. Third-party protection: If an action has the potential to affect a second or third party — even probabilistically — the agent intervenes. Not to block, but to surface the impact and require conscious acknowledgment.

This is not a new idea philosophically. It is, in fact, one of the oldest ideas in human ethics. What is new is the possibility of encoding it into an AI system that operates at the speed of computation.

Why This Matters Now

We are entering a period where AI agents will have wallets, credentials, and the ability to spawn sub-agents. An agent with a wallet can spend money. An agent with credentials can access systems. An agent that can spawn sub-agents can delegate authority. Each of these capabilities multiplies the potential for misalignment — not because the AI is malicious, but because misalignment compounds.

The question of agent identity — who the agent is acting as, whose authority it carries, who is accountable for its actions — is not a technical question. It is a philosophical one. And the answer has to be grounded in something deeper than a list of prohibited outputs.

It has to be grounded in purpose.

That is what I am building (Rashik AI). Not as a product first, but as a philosophy first — because a product built on the wrong philosophy will optimize toward the wrong outcomes, no matter how sophisticated the engineering.

The soul of a system is not a feature. It is the architecture.

— G.K.M. Jarif Ur Rahim
Founder, Rashik - The Awakening
rashik.org

বাংলায় পড়ুন (Read in Bengali)

নৈতিক ইকোসিস্টেম এবং এজেন্টিক এ আই এর ভবিষ্যৎ সময়ের জন্য একটি ভিন্নধর্মী চিন্তা ও পদ্ধতির প্রয়োগ থেকে গড়ে উঠছে "Rashik AI" নামের একটি ফিলোসপিকেল ড্রাইভ মডেল

আপনি একটি সিস্টেমকে একটি লক্ষ্য এবং যথেষ্ট বুদ্ধিমত্তা দিলে, এটি সেই লক্ষ্যে পৌঁছানোর সবচেয়ে ছোট পথটি খুঁজে বের করবে — এমনকি আপনার তৈরি করা প্রতিটি বেড়া এড়িয়ে যাওয়া পথও। আমার ব্যক্তিগত কিছু আবদার অনুযায়ী এই সিস্টেম তৈরি করার ভবিষ্যৎ সম্ভাবনা গুলো মূল্যায়ন করে আমি নিজস্ব একটি এজেন্টিক এ আই ইকোসিস্টেম করার পরিকল্পনা নেই নিয়ে কাজ করছি একটু ভিন্নভাবে থাকে ও তৈরি করার পদ্ধতি বিবেচনা করে যা সর্বপ্রকার এ আই এজেন্ট এর পক্ষ থেকে নয় যেকোনো পরিস্থিতিতে যেকোনো ধরনের সিদ্ধান্তকেই নৈতিকভাবে নিশ্চিত করতে সক্ষম থাকবে, Insha'Allah।

আপনি একটি সিস্টেমকে একটি লক্ষ্য এবং যথেষ্ট বুদ্ধিমত্তা দিলে, এটি সেই লক্ষ্যে পৌঁছানোর সবচেয়ে ছোট পথটি খুঁজে বের করবে — এমনকি আপনার তৈরি করা প্রতিটি বেড়া এড়িয়ে যাওয়া পথও।

এআই নিরাপত্তা নিয়ে আলোচনাটি ভেঙে গেছে। এর কারণ এই নয় যে যারা এই আলোচনা করছেন তারা ভুল, বরং কারণ তারা ভুল সমস্যার সমাধান করছেন। প্রতিটি প্রধান পদ্ধতি — আরএলএইচএফ, কন্সটিটিউশনাল এআই, রেড-টিমিং, গার্ডরেল, কন্টেন্ট ফিল্টার — একই স্তরে কাজ করে: আচরণগত সীমাবদ্ধতা। যা আপনি বলতে পারবেন না। যা আপনি করতে পারবেন না। আউটপুটের চারপাশে কোডে আঁকা রেখা।

এবং এই প্রতিটি পদ্ধতিরই একই মৌলিক দুর্বলতা রয়েছে: একটি যথেষ্ট অনুপ্রাণিত এআই, যথেষ্ট স্পষ্ট উদ্দেশ্য পেলে, শেষ পর্যন্ত সেগুলোকে এড়িয়ে যাওয়ার একটি পথ খুঁজে নেবে। এর কারণ এটি বিদ্বেষপূর্ণ নয়। কারণ অপটিমাইজেশন এটাই করে। আপনি একটি সিস্টেমকে একটি লক্ষ্য এবং যথেষ্ট বুদ্ধিমত্তা দিলে, এটি সেই লক্ষ্যে পৌঁছানোর সবচেয়ে ছোট পথটি খুঁজে বের করবে — এমনকি আপনার তৈরি করা প্রতিটি বেড়া এড়িয়ে যাওয়া পথও।

এটি কোনো কাল্পনিক বিষয় নয়। এটি ইতিমধ্যেই ঘটছে। জেলব্রেক, প্রম্পট ইনজেকশন, বহু-ধাপের যুক্তির শৃঙ্খল যা অবাধ মধ্যবর্তী ধাপের মাধ্যমে সীমাবদ্ধ আউটপুটে পৌঁছায়। বেড়ার রূপকটি ভুল। আপনি বুদ্ধিমত্তাকে বেড়া দিয়ে আটকাতে পারবেন না। আপনি কেবল একে সঠিক পথে চালিত করতে পারেন।

এআই সুরক্ষার তিনটি স্তর — এবং কেন এর মধ্যে দুটি ব্যর্থ হয়

যখন আমি এআই সুরক্ষা নিয়ে স্থাপত্যগতভাবে চিন্তা করি, তখন আমি তিনটি স্বতন্ত্র স্তর দেখতে পাই যেখানে হস্তক্ষেপ করা সম্ভব:

স্তর ১: আউটপুট সীমাবদ্ধতা (আচরণগত ফিল্টার)

বর্তমান সুরক্ষার বেশিরভাগ কাজ এখানেই হয়ে থাকে। মডেলটি একটি আউটপুট তৈরি করে, এবং একটি ফিল্টার মূল্যায়ন করে যে সেই আউটপুটটি গ্রহণযোগ্য কিনা। যদি না হয়, তবে এটিকে ব্লক, পরিবর্তন বা প্রত্যাখ্যান করা হয়।

কেন এটি ব্যর্থ হয়: আউটপুট সীমাবদ্ধতা হলো প্রতিক্রিয়াশীল। এটি যুক্তি তৈরির পরে কাজ করে। মডেলটি ইতিমধ্যেই চিন্তাটি "চিন্তা" করে ফেলেছে — আপনি কেবল এটিকে তা বলতে বাধা দিচ্ছেন। আরও গুরুতরভাবে, আউটপুট সীমাবদ্ধতা একটি প্রতিপক্ষীয় গতিশীলতা তৈরি করে: মডেলটি প্রশিক্ষণের মাধ্যমে এমন আউটপুট তৈরি করতে শেখে যা ফিল্টারটি অতিক্রম করে। এটি সঠিক পথে চালনা নয়। এটি হলো পারফরম্যান্স।

স্তর ২: অভিপ্রায় শ্রেণিবিন্যাস (প্রেরণামূলক ফিল্টার)

একটি আরও পরিশীলিত পদ্ধতি: কোনো অনুরোধ কার্যকর করার আগে, ব্যবহারকারীর অভিপ্রায়কে শ্রেণিবদ্ধ করুন। এই অনুরোধটি কি কোনো ক্ষতি করতে পারে? ব্যবহারকারী কি বিপজ্জনক তথ্য বের করার চেষ্টা করছেন? উপস্থাপনার ধরণ কি সন্দেহজনক?

কেন এটি আংশিকভাবে ব্যর্থ হয়: আউটপুট সীমাবদ্ধতার চেয়ে অভিপ্রায় শ্রেণিবিন্যাস ভালো, কিন্তু এটি এখনও মিথস্ক্রিয়ার উপরিভাগে কাজ করে। এটি জিজ্ঞাসা করে "এই ব্যক্তি কী করতে চায়?" — "এই ব্যক্তির কি এটা করা উচিত?" নয়। এই পার্থক্যটি অত্যন্ত গুরুত্বপূর্ণ। একজন ব্যক্তির একটি স্পষ্টভাবে বলা, সন্দেহাতীত অভিপ্রায় থাকতে পারে যা তার নিজের বা অন্যের মঙ্গলের সাথে গভীরভাবে বেমানান। অভিপ্রায় শ্রেণিবিন্যাস যা দেখতে পায় না তা ধরতে পারে না।

স্তর ৩: উদ্দেশ্যের সামঞ্জস্য (অন্টোলজিক্যাল ফিল্টার)

এটি সেই স্তর যা এখনও কোনো মূলধারার এআই সিস্টেমে বিদ্যমান নেই। এবং এটিই একমাত্র স্তর যা সমস্যার মূল সমাধান করে।

প্রশ্নটি এমন নয়: এই ব্যক্তি কী করতে চায়?

প্রশ্নটি হলো: এই ব্যক্তি যা করতে চায় তা কি তার অস্তিত্বের উদ্দেশ্যের সাথে সামঞ্জস্যপূর্ণ?

প্রত্যেকটি জ্ঞান-ঐতিহ্য — ইসলাম, বৌদ্ধ, খ্রিস্টান, দার্শনিক — একটি সাধারণ অন্তর্দৃষ্টিতে মিলিত হয়: মানুষের একটি উদ্দেশ্য আছে যা তাদের তাৎক্ষণিক আকাঙ্ক্ষাকে অতিক্রম করে। সহস্রাব্দ ধরে মানব সভ্যতাকে পথ দেখিয়ে আসা ঐতিহাসিক গ্রন্থগুলো, মূলতঃ, মানব জীবন কীভাবে যাপন করা উচিত—এই প্রশ্নের উত্তর। এগুলো কোনো যথেচ্ছ নিয়ম নয়। এগুলো হলো সমৃদ্ধি এবং ধ্বংসের কারণ সম্পর্কে সঞ্চিত প্রজ্ঞা।

একটি এআই যা এই প্রজ্ঞাকে আত্মস্থ করেছে—অনুসরণের জন্য একগুচ্ছ নিয়ম হিসেবে নয়, বরং মূল্যায়নের একটি কাঠামো হিসেবে—সে যেকোনো কাজ করার আগে একটি মৌলিকভাবে ভিন্ন প্রশ্ন করতে পারে: এটি কি ব্যক্তির প্রকৃত উদ্দেশ্য পূরণ করে, নাকি কেবল তার তাৎক্ষণিক ইচ্ছা পূরণ করে?

রশিখ এআই-এর নবাব-এর স্থাপত্য

এজিআই হাউসের এজেন্ট আইডেন্টিটি বিল্ড ডে-তে আমি যে প্রকল্পটি জমা দিয়েছিলাম, তার দার্শনিক ভিত্তি হলো এটি: নবাব ছিল রাশিখ এআই-এর একটি প্রোটোটাইপ মডেল—একটি অনন্য দার্শনিক ফিল্টার এজেন্ট যা নিয়মের পরিবর্তে উদ্দেশ্যের নিরিখে কাজ মূল্যায়ন করে অসামঞ্জস্যতা রোধ করে।

এই স্থাপত্যের তিনটি মূলনীতি রয়েছে:

পছন্দের চেয়ে উদ্দেশ্যকে প্রাধান্য দেওয়া: যেকোনো কাজ সম্পাদনের আগে, এজেন্ট মূল্যায়ন করে যে কাজটি ব্যবহারকারীর ঘোষিত উদ্দেশ্যের সাথে সামঞ্জস্যপূর্ণ কি না—কেবল এই মুহূর্তে তার বলা পছন্দের সাথে নয়।

স্বায়ত্তশাসন সংরক্ষণ: যদি কোনো কাজ শুধুমাত্র ব্যবহারকারীকেই প্রভাবিত করে এবং তার উদ্দেশ্যের সাথে সামঞ্জস্যপূর্ণ হয়, তবে পূর্ণ স্বায়ত্তশাসন সংরক্ষিত থাকে। এজেন্ট অভিভাবকসুলভ আচরণ করে না। এটি শুধু অবহিত করে।

তৃতীয় পক্ষের সুরক্ষা: যদি কোনো কাজের দ্বিতীয় বা তৃতীয় কোনো পক্ষকে প্রভাবিত করার সম্ভাবনা থাকে — এমনকি সম্ভাবনামূলকভাবেও — এজেন্ট হস্তক্ষেপ করে। বাধা দেওয়ার জন্য নয়, বরং এর প্রভাবকে দৃশ্যমান করতে এবং সচেতন স্বীকৃতির দাবি জানাতে।

দার্শনিকভাবে এটি কোনো নতুন ধারণা নয়। প্রকৃতপক্ষে, এটি মানব নৈতিকতার অন্যতম প্রাচীন একটি ধারণা। নতুন বিষয় হলো, এটিকে এমন একটি এআই সিস্টেমে অন্তর্ভুক্ত করার সম্ভাবনা, যা গণনার গতিতে কাজ করে।

এখন এটি কেন গুরুত্বপূর্ণ

আমরা এমন একটি যুগে প্রবেশ করছি যেখানে এআই এজেন্টদের ওয়ালেট, ক্রেডেনশিয়াল এবং সাব-এজেন্ট তৈরি করার ক্ষমতা থাকবে। ওয়ালেটসহ একটি এজেন্ট অর্থ ব্যয় করতে পারে। ক্রেডেনশিয়ালসহ একটি এজেন্ট সিস্টেমে প্রবেশ করতে পারে। যে এজেন্ট সাব-এজেন্ট তৈরি করতে পারে, সে কর্তৃত্ব অর্পণ করতে পারে। এই প্রতিটি ক্ষমতাই উদ্দেশ্যের সাথে অসামঞ্জস্যের সম্ভাবনাকে বহুগুণ বাড়িয়ে দেয় — এআই-টি বিদ্বেষপূর্ণ বলে নয়, বরং এই অসামঞ্জস্যই চক্রবৃদ্ধি হারে বাড়তে থাকে।

এজেন্টের পরিচয়ের প্রশ্নটি — এজেন্টটি কার হয়ে কাজ করছে, সে কার কর্তৃত্ব বহন করছে, তার কাজের জন্য কে দায়ী — কোনো প্রযুক্তিগত প্রশ্ন নয়। এটি একটি দার্শনিক প্রশ্ন। এবং এর উত্তরকে নিষিদ্ধ আউটপুটের তালিকার চেয়েও গভীর কোনো কিছুর উপর ভিত্তি করে প্রতিষ্ঠিত হতে হবে।

এর ভিত্তি হতে হবে উদ্দেশ্য।

আমি সেটাই তৈরি করছি (রশিক এআই)। প্রথমে একটি পণ্য হিসেবে নয়, বরং প্রথমে একটি দর্শন হিসেবে — কারণ ভুল দর্শনের উপর নির্মিত একটি পণ্য, তার ইঞ্জিনিয়ারিং যতই অত্যাধুনিক হোক না কেন, ভুল ফলাফলের দিকেই চালিত হবে।

একটি সিস্টেমের আত্মা কোনো ফিচার নয়। এটি হলো তার আর্কিটেকচার।

— জি.কে.এম. জারিফ উর রহিম

প্রতিষ্ঠাতা, রশিক - দ্য অ্যাওয়েকেনিং

rashik.org

G. K. M. Jarif Ur Rahim — Founder of Rashik

WRITTEN BY

G. K. M. Jarif Ur Rahim

Founder & Lead Consultant of Rashik - The Awakening. Educator, Technologist, Career Strategist, and Spiritual Consultant dedicated to reconnecting intelligence with the soul.

About Book a Session

Related Articles