رجوع للمدوّنة

Anthropic فتحت باب الإبلاغ عن ثغرات jailbreak بـ Claude Fable 5

إطار عمل لتقييم خطورة الـ jailbreaks مع قناة إبلاغ عامة عبر HackerOne لـ Fable 5. الجزء اللي بيفيدك كباحث أو مستخدم هو إنو صار في باب رسمي للإبلاغ.

شو اللي صار يوم 2 تمّوز

Anthropic نزّلت يوم 2 تمّوز تفصيل جديد عن حمايات Fable 5 الإلكترونية، ومعه مسوّدة أوّلية لإطار عمل (framework) لقياس خطورة محاولات كسر الحمايات (jailbreaks)، طوّروه مع شركاء Project Glasswing. بالمرة، فتحوا برنامج إبلاغ عام على HackerOne خاص بـ cyber jailbreaks — يعني أي باحث أمني يقدر يرسل اكتشافه بشكل منظّم.

هاد مش خبر تقني جاف. هاد بيتكلم عن شي بيمسّك كل شخص عم يستخدم Claude Fable 5: شو بيصير إذا واحد لقى طريقة يخلّي النموذج يعمل شي ما مفروض يعمله، وكيف Anthropic عم تحاول ترتّب الفوضى قبل ما تتفاقم.

شو يعني "jailbreak" بالضبط

ببساطة، الـ jailbreak هو أي طريقة بتخلّي النموذج يتجاوز قواعده. مش بس "اسأله شو ما بدك" — في تقنيات متطوّرة بتستخدم قصص معقّدة، سيناريوهات لعب أدوار، أو حتّى payloads مكتوبة بالـ code.

Fable 5 هاد أول نموذج من Anthropic يجي بـ safety classifiers متقدّمة — يعني "أربع فئات" للسلوك:

  • ممنوع (prohibited): طلبات متل مساعدة بكتابة malware أو عمليات اختراق. النظام بيرفض مباشرة.
  • ثنائي الاستخدام عالي الخطورة (high-risk dual-use): طلبات ممكن تنفّذ بشكل مشروع أو ضار. هون في نقاش.
  • ثنائي الاستخدام منخفض الخطورة (low-risk dual-use): أدوات فيها احتمال خطر بس واقعياً نادر.
  • آمن (benign): الطلب العادي.

اللي صار قبل هيك: النموذج كان يرفض أو يقبل بدون تفسير. هلّق، إذا تم الرفض، الـ API بترجع `stop_reason: "refusal"` كاستجابة ناجحة (HTTP 200) مع اسم الـ classifier اللي رفض. يعني المطوّر بيعرف ليش تم الرفض، والـ enterprise بتقدر تاخذ قرارها بنفسها.

ليش هاد الـ framework مهم

قبل هيك، كل مختبر كان عم يقيس خطورة الـ jailbreaks بطريقته. واحد بيقول "كارثة"، وآخر بيقول "بسيط". ما كان في مقياس موحّد. النتيجة: النقاش العام عن "هل AI خطير؟" كان قائم على انطباعات.

الـ framework الجديد بيحاول يحل هاد بـ:

  1. مقياس رقمي موحّد للخطورة — أي باحث، أي مختبر، أي حكومة بتقدر تقرأ نفس الرقم.
  2. تصنيف واضح للـ payload — مش بس "نجح" أو "ما نجح".
  3. قناة إبلاغ رسمية (HackerOne) — أي اكتشاف بيتسجّل، بيتدرّج، بيتنشر للناس اللي يهمّها.

بالمختصر: بدل ما السوق يحكم بانطباعات، صار في مقياس فعلي. هاد بيسهّل النقاش عن قيود التصدير، وعن متى يجب على المختبر يسحب نموذج من السوق، وكمان بيسهّل على الباحثين الشغل بدون ما يخافوا من ملاحقة قانونية.

شو الجديد عملياً للباحثين

قبل هيك، إذا باحث لقى jailbreak بـ Claude، كان عنده خيارين سيّئين:

  • ينشره علناً — فيخلي المستخدمين العاديين يستفيدوا، بس بنفس الوقت يفتح الباب للمشاكل.
  • يبعت لـ Anthropic بس — فيأخذ وقت طويل، وما حدا بيعرف شو صار.

البرنامج الجديد على HackerOne بيعطي:

  • مسار إبلاغ منظّم بـ scope واضح.
  • اعتراف علني (public disclosure) بعد ما يصير في إصلاح.
  • مكافآت (bounties) للأبحاث عالية الجودة — حسب شو مكتوب بـ scope.

إذا كنت باحث أمني أو مهتم بـ AI safety، هاد باب رسمي صار مفتوح. ما بقا مضطر تتواصل بشكل غير رسمي مع الشركة أو تخاف من الملاحقة.

⚡ جرّبها بنفسك

إذا كنت باحث أو مستخدم متقدّم:

  1. افتح صفحة Anthropic على HackerOne وتفقّد الـ scope. شو الأنواع المقبولة للإبلاغ، وشو مش مقبول.
  2. إذا عندك اختبار آمن لـ Fable 5 (مش على production) وجربت prompt يعمل refusal، سجّل الـ prompt والـ response بالتفصيل.
  3. قبل ما تبلّغ، اقرأ شروط البرنامج: شو بيصير بالاعتراف العلني، متى رح تتواصل معك، وهل في مكافأة.

إذا كنت مستخدم عادي:

  • الخبر هاد ما بيغيّر شو بتعمل بـ Claude — استخدمه متل قبل.
  • بس إذا شفت بأخبار تالية "Anthropic أوقفت jailbreak جديد"، هلّق بتعرف هاد جزء من برنامج منظّم، مش فوضى.

خلاصة بجملة وحدة

لأول مرة، عندنا مقياس موحّد + باب إبلاغ رسمي لمحاولات كسر حمايات نماذج الـ AI الحدودية. هاد بيخلّي النقاش العام عن أمان الـ AI قائم على أرقام مش انطباعات، وبيعطي الباحثين مسار واضح بدل الفوضى اللي كانت قبل.

المصادر

النسخة الإنجليزية

Anthropic opens a HackerOne program for Claude Fable 5 jailbreaks

A jailbreak-severity scoring framework plus a public HackerOne disclosure program for Fable 5. The part that matters for everyday users is the new public reporting channel.

What happened on July 2

Anthropic published a detailed breakdown of the safety classifiers behind Fable 5 on July 2, 2026, alongside an early draft of a jailbreak-severity scoring framework developed with Project Glasswing partners. They also opened a public HackerOne disclosure program specifically for cyber jailbreaks on Fable 5.

This is not just a technical post. It speaks to anyone using Claude Fable 5: what happens when someone finds a way to bypass the model's safety rules, and how Anthropic is trying to bring order before things get out of hand.

What "jailbreak" actually means here

A jailbreak is any technique that gets a model to bypass its safety rules. Not just "ask it anything" — there are sophisticated methods using elaborate scenarios, role-play setups, or code-formatted payloads.

Fable 5 is the first Anthropic model shipped with advanced safety classifiers that bucket behavior into four categories:

  • Prohibited: requests like writing malware or assisting with exploits. The system refuses directly.
  • High-risk dual-use: requests that could be legitimate or harmful. This is where the interesting work happens.
  • Low-risk dual-use: tools with some risk but rare real-world abuse.
  • Benign: ordinary requests.

Previously, the model would just refuse or accept without explanation. Now, when a refusal happens, the API returns `stop_reason: "refusal"` as a successful HTTP 200 response, with the name of the classifier that triggered the refusal. Developers know exactly why, and enterprise customers can make their own call.

Why the framework matters

Until now, every lab measured jailbreak severity its own way. One would call it "catastrophic", another "minor". There was no shared scale. The public conversation about "is AI dangerous?" was running on vibes.

The new framework tries to fix this with three things:

  1. A unified numeric severity scale that any researcher, lab, or government can read the same way.
  2. A clear classification of the payload itself — not just "worked" or "did not work".
  3. An official disclosure channel (HackerOne) where findings are recorded, triaged, and shared with the people who need to know.

In short: instead of the market judging on impressions, there is now a real yardstick. That makes the export-control debate easier, makes it clearer when a lab should pull a model, and lets researchers work without legal risk.

What changes practically for researchers

Before this, a researcher who found a Claude jailbreak had two bad options:

  • Publish it publicly — making it accessible to regular users but also opening the door to problems.
  • Send it to Anthropic only — slow, and nobody else ever knows what happened.

The new HackerOne program offers:

  • A structured disclosure path with a clear scope.
  • Public disclosure once a fix is in place.
  • Bounties for high-quality findings — depending on the published scope.

If you are a security researcher or care about AI safety, this door is officially open. You no longer have to reach out informally or worry about legal action.

⚡ Try it yourself

If you are a researcher or advanced user:

  1. Open Anthropic's HackerOne page and review the scope. What types of reports are accepted, what is not.
  2. If you have a safe test environment for Fable 5 (not production) and you find a prompt that triggers a refusal, document the prompt and the response in detail.
  3. Before reporting, read the program terms: what happens with public disclosure, when they will reach out, and whether a bounty applies.

If you are a regular user:

  • This news does not change how you use Claude — keep using it as before.
  • But if you see follow-up stories about "Anthropic blocked a new jailbreak", you now know that is part of an organized program, not chaos.

One-line takeaway

For the first time, we have a shared severity scale and an official disclosure channel for jailbreak attempts on frontier AI models. The public conversation about AI safety is now running on numbers, not impressions, and researchers have a clear path forward.

رجوع لكل المقالات