رجوع للمدوّنة

نماذج Claude الجديدة بتستدعي الأدوات بشكل أضعف

مبتكر Flask بيقول إن Opus 4.8 وSonnet 5 صاروا يطلّعوا استدعاءات أدوات معطوبة أكتر من النماذج الأقدم. إذا بتبني وكلاء، هاد بيضرك.

شو لقى Armin Ronacher فعلياً

Armin Ronacher مش مستخدم AI عادي. هو مبتكر Flask وSentry، وواحد من أكتر باني الأدوات المحترمين بعالم Python. لمّا بيكتب عن تراجع، يستاهل القراءة.

يوم 4 تموز 2026، نشر بوست عنوانه "Better Models: Worse Tools". الفكرة الرئيسية: نماذج Claude الأحدث — Opus 4.8 وSonnet 5 — صاروا يطلّعوا استدعاءات أدوات معطوبة أكتر من النماذج الأقدم. مش أقل. أكتر. وبيدعم هاد بلوغات إنتاج من أنظمة تبعه هو.

النمط بيبيّن بوضوح. كل ما كان النموذج أحدث، كل ما صار يخترع حقول ما إياها بالمخطط، يبعت JSON ناقص ما بيقبل parse، أو يلف استدعاء الأداة بصياغة مختلفة شوي عن المتوقع. سمّاه "تراجع مُضايق" — وأي حدا بيشغّل loop وكيل رح يتعرّف ع هالإحساس.

ليش هاد بيهمك إذا بتبني وكلاء

إذا بتكتب كود بينادي Claude عبر API وبيرجّع نتائج الأدوات، كل استدعاء معطوب هو فشل لازم تتعامل معه. النموذج مو دايماً صح لأنه أحدث. دقة استدعاء الأدوات محور منفصل عن جودة الاستدلال، وع هالمحور، النماذج الأحدث عم تسوء.

لـ chatbot هاد بيكون مخفي. تسأل سؤال، بتاخد جواب، بتكمّل تتصفّح. لـ loop وكيل بينادي الأدوات عشرات المرات بكل مهمة، التراجع بيبيّن بكل مكان: أكتر محاولات إعادة، أكتر أخطاء تحقق، أكتر توكنات ضايعة، أكتر وقت محسوب ع شغل النموذج عمل نصّو. المستخدم بيحسّها متل "الوكيل اليوم مش مستقر".

لشركة عم تطلق منتجات AI، الوضع أسوأ. معدّل الخطأ ما بينزل لمّا بترقّي للنموذج الأحدث — بيطلع. وعد "الأحدث = الأفضل" بينكسر بهدوء بلحظة ما يكون عندك مخططات وكلاؤك معتمدين إلها.

شو بيفسّر Ronacher (الجزء التقني)

أهم قسم بالبوست هو التفسير التقني. استدعاءات الأدوات بـ Claude مو قناة بروتوكول منفصلة — هي نصّ عادي داخل الرد مع علامات ANTML. النموذج بيكتب لغة طبيعية، بيضرب علامة، بيكتب JSON، بيضرب علامة ثانية. إذا النموذج بيعرف المخطط منيح، بيطلع JSON نظيف. إذا المخطط غريب أو غير مألوف، النموذج لازم يجتهد — وجودة هالاجتهاد بتعتمد ع التدريب، مو ع الذكاء الخام.

Ronacher بيقول إن نماذج Claude الأحدث يبدو إنها صرفت وقت تدريب أكتر ع الاستدلال طويل الشكل وأقل ع انضباط المخرجات المنظمة. النتيجة: مقالات أحسن، انضباط مخططات أضعف. خسارة صافية لأي حدا بيبني وكلاء.

وكمان بيعرض أمثلة ملموسة. بمثال واحد، Opus 4.8 بينادي أداة بكائن JSON فيه اسم حقل مختلف شوي عن المخطط الموثّق — قريب للدرجة يلي بيبيّن صحيح، بس مختلف للدرجة يلي بيكسر المحقق. بمثال ثاني، Sonnet 5 بيحذف حقل مطلوب بالكامل، وكودك ما بيلحق يكتشف الخطأ إلا بعد ما استدعاء الدالة بلّش ينفّذ.

شو يعني هاد إذا إنت مش مطوّر

حتى لو ما بتكتب وكيل أبداً، هاد بيأثّر عليك. نفس النماذج بتشغّل أغلب أدوات AI يلي بتستخدمها — ملحقات المتصفح، مساعدات البريد، مساعدات الجداول، مساعدات البرمجة. إذا هالأدوات عندها تكلفة مخفية من استدعاءات معطوبة، هالتكلفة بتبيّن متل ردود أبطأ، خلل متفرّق، أو ميزات بتشتغل نهار وبتنكسر نهار ثاني.

وفي نقطة أعمق عن الحوافز. مختبرات AI عم تتسابق تطلق نماذج "أذكى" ع المعايير. بس المعايير ما بتقيس انضباط المخططات. بتقيس الاستدلال. يعني نموذج يقدر يصعّد ليدربورد وبنفس الوقت يسوء بالشغل المنظّم يلي التطبيقات الحقيقية معتمدة عليه. المستخدمين بيدفعوا ع البريق، الشركات بتدفع ع الاحتكاك.

⚡ شو تعمل هلأ

ما لازم تستنى Anthropic تطلق إصلاح. في كم خطوة عملية بتساعد فوراً:

  • ضيف محقق JSON صارم قبل ما تبعت استدعاءات الأدوات لكودك. ارفض أي شي ما بيتماشى مع المخطط، واطلب من النموذج يعيد المحاولة مع تذكير صريح بالمخطط.
  • احتفظ بقائمة "أنماط معطوبة معروفة" وتحقق منها قبل التنفيذ. هاد بيمسك الأنماط الشائعة قبل ما توصل لـ function عندك.
  • ثبّت نسخة النموذج حسب الاستخدام. استخدم Sonnet 5 للدردشة والتلخيص، بس خلّي نموذج Claude أقدم (أو عائلة مختلفة) لـ loops الوكلاء الكتيرة بالأدوات، لحتى يتحسّن التراجع.
  • سجّل كل فشل بأداة مع اسم النموذج والنص المعطوب. بعد أسبوع رح تشوف أي نموذج فعلاً الأنسب لـ stack تبعك — وأي نموذج الأكتر تكلفة للتشغيل.
  • إذا عم تختار نموذج لمنتج جديد، اعمل benchmark صغير للوكيل قبل ما تلتزم. عشرين دقيقة اختبار بتوفّرك شهرين ديباغ بالإنتاج.

🔧 جرّبها بنفسك

اختار وكيل أو workflow بيستخدم أدوات عم تشغّله اليوم. بدّل نموذجه لأي شي عندك لـ Opus 4.8 أو Sonnet 5 لساعة وحدة. عدّ الاستدعاءات المعطوبة بالمهمة. من بعدها ارجّع للنموذج الأصلي. إذا لاحظت قفزة واضحة بالأخطاء، عندك نقطة بيانات خاصة فيك — وبتقدر تقرر أي نسخة نموذج بأي جزء من نظامك. تجربة بعشر دقائق ممكن توفّرلك ربع ديباغ.

📚 المصادر

  1. Armin Ronacher — "Better Models: Worse Tools"
  2. توثيق استدعاء الأدوات بـ Anthropic
  3. نقاش Hacker News عن البوست

المصادر

النسخة الإنجليزية

Newer Claude models call tools worse — Armin Ronacher has the receipts

The creator of Flask says Opus 4.8 and Sonnet 5 emit more malformed tool calls than older Claude models. If you build agents, this hurts directly.

What Armin Ronacher actually found

Armin Ronacher is not a casual AI user. He's the creator of Flask and Sentry, and one of the most respected tool-builders in the Python world. When he writes about a regression, it's worth reading.

On July 4, 2026, he published a post titled "Better Models: Worse Tools." The thesis: newer Claude models — Opus 4.8 and Sonnet 5 — emit MORE malformed tool calls than older ones. Not fewer. More. He backs it up with production logs from his own systems.

The pattern shows up clearly. The newer the model, the more often it invents schema fields that don't exist, sends partial JSON that won't parse, or wraps tool calls in slightly the wrong syntax. He calls it an "aggravating regression" — and anyone running an agent loop will recognize the feeling.

Why this matters if you build agents

If you write code that calls Claude through an API and feeds tool results back, every malformed tool call is a failure you have to handle. The model is not always right just because it's newer. Tool-call fidelity is a separate axis from reasoning quality, and on that axis the newer models are doing worse.

For a chatbot this is invisible. You ask a question, you get an answer, you scroll on. For an agent loop that calls tools dozens of times per task, the regression shows up everywhere: more retries, more validator errors, more wasted tokens, more time billed for work the model half-did. The user feels it as "the agent is flaky today."

For a company shipping AI products, this is even worse. Your error rate doesn't go down just because you upgraded to the latest model — it goes up. The promise of "newer = better" quietly breaks the moment you have schemas your agents depend on.

What Ronacher actually explains (the technical bit)

The most useful section of the post is the technical explanation. Tool calls in Claude are not a separate protocol channel — they are in-band text with ANTML markers. The model writes natural language, hits a marker, writes JSON, hits another marker. If the model knows the schema well, it produces clean JSON. If the schema is unusual or unfamiliar, the model has to improvise — and the quality of that improvisation depends on training, not raw intelligence.

Ronacher's argument is that newer Claude models appear to have spent more training time on long-form reasoning and less on disciplined structured output. The result: better essays, worse schema discipline. A net loss for anyone building agents.

He also shows concrete examples. In one case, Opus 4.8 calls a tool with a JSON object that contains a field name slightly different from the documented schema — close enough to look right, different enough to break the validator. In another, Sonnet 5 drops a required field entirely and the user's code only catches the error after the function call has already started executing.

What this means if you're not a developer

Even if you never write an agent yourself, this affects you. The same models power most of the AI tools you use — browser extensions, email assistants, spreadsheet helpers, coding copilots. If those tools have a hidden cost from malformed tool calls, that cost shows up as slower responses, occasional glitches, or features that work one day and break the next.

There's also a deeper point about incentives. AI labs are racing to ship "smarter" models on benchmarks. But benchmarks don't measure schema discipline. They measure reasoning. So a model can climb the leaderboards while quietly getting worse at the unglamorous structured-output work that real applications depend on. Users pay for the gloss, businesses pay for the friction.

⚡ What to do about it today

You don't need to wait for Anthropic to ship a fix. A few practical moves help immediately:

  • Add a strict JSON validator before sending tool calls to your code. Reject anything that doesn't parse or doesn't match your schema, and ask the model to retry with an explicit reminder of the schema.
  • Keep a list of "known bad" patterns and pre-validate against them. This catches the common failure modes before they hit your function.
  • Pin your model version per use case. Use Sonnet 5 for chat and summarisation, but keep an older Claude (or a different family) for tool-heavy agent loops until the regression improves.
  • Log every tool-call failure with the model name and the malformed output. After a week you'll see which model is actually the best fit for your stack — and which one is the most expensive to operate.
  • If you're choosing a model for a new product, run a small agent benchmark before committing. Twenty minutes of testing beats two months of debugging in production.

🔧 Try it yourself

Pick one agent or tool-using workflow you run today. Switch its model from whatever you have to Opus 4.8 or Sonnet 5 for one hour. Count the failed tool calls per task. Then switch back. If you see a clear jump in errors, you have your own data point — and you can decide which model version each part of your system should use. That's a ten-minute experiment that can save you a quarter of debugging.

📚 Sources

  1. Armin Ronacher — "Better Models: Worse Tools"
  2. Anthropic tool-use documentation
  3. Hacker News discussion of the post
رجوع لكل المقالات