GitHub Actions merge queue के अंदर आठ महीने बिताने से आप कुछ ऐसा सीखते हैं जो फीचर तुलना मैट्रिक्स (feature comparison matrices) कभी नहीं सिखा सकते। एक फ्रेमवर्क पचास मेट्रिक्स, शानदार डैशबोर्ड और प्रतिष्ठित रिसर्च लैब्स के उद्धरण (citations) पेश कर सकता है। लेकिन अगर वह आपके डिप्लॉयमेंट को इसलिए रोकता है क्योंकि एक "vibe check" स्कोर समान कोड के लिए 0.72 से गिरकर 0.68 हो गया है, तो वह उपयोगी होने के बजाय नुकसानदेह है। यह आपकी शिपिंग वेलोसिटी (shipping velocity) के लिए एक सक्रिय खतरा बन जाता है।
यही वह फिल्टर है जिसे अधिकांश LLM इवैल्यूएशन राउंडअप (evaluation roundups) मिस कर देते हैं। वे क्षमताओं (capabilities) को गिनते हैं। वे शायद ही कभी वह एकमात्र सवाल पूछते हैं जो एक मर्ज क्यू (merge queue) में मायने रखता है: क्या यह चेक हर बार चलने पर बिल्कुल उसी तरह पास और फेल होता है?
मैंने यह कठिन काम करके सीखा। मैंने छह ओपन-सोर्स LLM इवैल्यूएशन फ्रेमवर्क्स को एक वास्तविक CI पाइपलाइन से जोड़ा। वे आठ महीनों तक लाइव प्रोडक्शन पुल रिक्वेस्ट (pull requests) पर चले। दो ने गेटकीपर (gatekeepers) बने रहने का अधिकार कमाया। बाकी को एडवाइजरी डैशबोर्ड (advisory dashboards) में बदल दिया गया, नाइटली जॉब्स (nightly jobs) में डाल दिया गया, या पूरी तरह से हटा दिया गया। सबक कड़ा और महंगा था: जब आप मेन ब्रांच (main branch) की रक्षा कर रहे हों, तो प्रोबेबिलिस्टिक क्वालिटी (probabilistic quality) के बजाय डिटरमिनिस्टिक स्ट्रक्चर (deterministic structure) काम आता है।
एक मर्ज गेट (Merge Gate) का असली काम
एक CI गेट रिसर्च एनवायरनमेंट नहीं है। यह एक बाउंसर है। इसका पूरा उद्देश्य एक विशिष्ट बदलाव को देखना और 'हाँ' या 'ना' में उत्तर देना है। हाँ, यह PR मेन ब्रांच में शामिल हो सकता है। नहीं, यह नहीं हो सकता। वह उत्तर सेकंडों में आना चाहिए, बहुत कम लागत वाला होना चाहिए, और कभी भी पिछले परिणामों को नहीं बदलना चाहिए। यदि आप एक शांत मंगलवार और एक भागदौड़ भरे शुक्रवार को उसी कमिट (commit) के खिलाफ उसी पाइपलाइन को फिर से चलाते हैं, तो परिणाम बिल्कुल समान होना चाहिए।
यहीं पर अधिकांश LLM इवैल्यूएशन फ्रेमवर्क लड़खड़ा जाते हैं। उन्हें डेटा वैज्ञानिकों द्वारा डेटा वैज्ञानिकों के लिए बनाया गया है। वे इनसाइट (insight), एक्सप्लोरेशन (exploration) और सूक्ष्म स्कोरिंग (nuanced scoring) के लिए ऑप्टिमाइज़ किए गए हैं। जबकि एक मर्ज क्यू बाइनरी निर्णयों (binary decisions), गति और जीरो फ्लैकीनेस (zero flakiness) के लिए ऑप्टिमाइज़ होता है। ये दोनों लक्ष्य केवल आंशिक रूप से ही मिलते हैं।
LLM-as-Judge मर्ज क्यू को क्यों बिगाड़ देता है
मेरे टेस्ट में विफल रहे टूल्स में एक ही डिज़ाइन दोष (design sin) था: वे प्राथमिक गेट मैकेनिज्म के रूप में LLM-as-judge कॉल्स पर बहुत अधिक निर्भर थे।
एक LLM-as-judge प्रॉम्प्ट मॉडल से आउटपुट को एक से दस के पैमाने पर स्कोर करने, या दो प्रतिक्रियाओं में से बेहतर को चुनने, या तथ्यात्मक सटीकता को रेट करने के लिए कहता है। क्वालिटी ट्रेंड्स को समझने के लिए यह दृष्टिकोण शक्तिशाली है। लेकिन एक ब्लॉकिंग CI चेक के लिए यह जहर के समान है। एक ही इनपुट अलग-अलग दिनों में अलग-अलग स्कोर दे सकता है क्योंकि टेम्परेचर (temperature), मॉडल वर्जनिंग और प्रॉम्प्ट फॉर्मेटिंग सभी शोर (noise) पैदा करते हैं। जब वह स्कोर एक हार्ड थ्रेशोल्ड (hard threshold) और हार्ड एग्जिट कोड (hard exit code) से जुड़ा होता है, तो आपकी क्यू 'भूतों' (ghosts) के कारण रुक जाती है।
विफलताएं तेजी से बढ़ती हैं। एक नॉन-डिटरमिनिस्टिक (nondeterministic) चेक क्यू बैकअप पैदा करता है। इंजीनियर तब तक दोबारा प्रयास (retry) करना सीख जाते हैं जब तक कि नंबर अनुकूल न आ जाए, जो टीम को रेड बिल्ड्स (red builds) को नजरअंदाज करने के लिए प्रशिक्षित करता है। टोकन लागत बढ़ जाती है क्योंकि हर रिट्राय में अधिक API क्रेडिट खर्च होते हैं। सबसे बुरा यह है कि सिग्नल अर्थहीन हो जाता है। एक रेड बिल्ड का मतलब होना चाहिए "आपने एक बग पेश किया है।" यदि इसका मतलब है "जज मॉडल आज थोड़ा नखरेबाज हो गया है," तो भरोसा खत्म हो जाता है।
जीवित बचे फ्रेमवर्क्स क्या अलग करते हैं
Promptfoo और DeepEval इसलिए बच गए क्योंकि वे डिटरमिनिस्टिक चेक को 'फर्स्ट-क्लास सिटीजन' मानते हैं और LLM जज स्कोर को सेकेंडरी, नॉन-ब्लॉकिंग सिग्नल के रूप में देखते हैं। वे समझते हैं कि एक गेट को एग्जिट कोड की आवश्यकता होती है, न कि किसी राय वाले फ्लोटिंग-पॉइंट नंबर की।
Promptfoo, जिसे MIT लाइसेंस के तहत जारी किया गया है, कमांड लाइन के लिए बनाया गया है। यह regex matches, JSON schema validation, contains checks और exact string comparisons जैसे एसर्शन (assertions) चलाता है। ये कोई फैंसी चीज़ें नहीं हैं। ये उन्नत grep और jq कमांड्स की तरह हैं। यही कारण है कि वे CI में काम करते हैं। एक regex या तो मैच करता है या नहीं करता। एक JSON schema या तो वैलिडेट करता है या एरर देता है। Promptfoo स्टैंडर्ड Unix exit codes लौटाता है, इसलिए GitHub Actions स्वाभाविक रूप से समझ जाता है कि मर्ज कब रोकना है। यह लैंग्वेज-एग्नोस्टिक (language-agnostic) है क्योंकि यह एक CLI टूल के रूप में काम करता है। आपको केवल आउटपुट को वैलिडेट करने के लिए Node.js सर्विस रेपो के अंदर पायथन इकोसिस्टम इंस्टॉल करने की आवश्यकता नहीं है।
DeepEval, जो Apache 2.0 के तहत लाइसेंस प्राप्त है, पायथन टीमों के लिए बेहतरीन विकल्प है। यह pytest की तरह इंटीग्रेट होता है। आप परिचित सिंटैक्स में टेस्ट लिखते हैं, और विफलता स्वाभाविक रूप से सूट को ब्लॉक कर देती है। DeepEval मेट्रिक्स का एक विशाल कैटलॉग प्रदान करता है, लेकिन महत्वपूर्ण बात यह है कि आपको उनका सावधानीपूर्वक उपयोग करना चाहिए। गेट्स के लिए डिटरमिनिस्टिक या ह्यूरिस्टिक (heuristic) मेट्रिक्स पर भरोसा करें। यदि आप G-Eval या अन्य जज-आधारित स्कोरर्स का उपयोग करते हैं, तो उन्हें हार्ड एसेर्ट्स (hard asserts) के बजाय नॉन-ब्लॉकिंग रिपोर्ट जनरेटर में लपेटें। इस तरह उपयोग करने पर, DeepEval आपको रिसर्च नोटबुक की अनिश्चितता के बिना एक टेस्टिंग फ्रेमवर्क की सुगमता (ergonomics) प्रदान करता है।
बाकी चार कहाँ फिट होते हैं
वे चार फ्रेमवर्क जो गेट के रूप में जीवित नहीं रह सके, उनका मूल्य अभी भी है। वे बस आपके टूलचेन (toolchain) में कहीं और के लिए हैं।
Future AGI (Apache 2.0) ships over fifty metrics and targets teams building custom SDKs. The metrics are thorough. The problem is that the tool expects you to write your own harness to drive it in a CI queue. In a research context, that is a reasonable trade. In a merge queue, every layer of custom wiring is a new source of instability. It is a capable evaluation engine, but not a ready gatekeeper.
RAGAS (Apache 2.0) excels at measuring retrieval-augmented generation quality. Its faithfulness and answer relevance metrics are genuinely useful for understanding how a knowledge base performs over time. Unfortunately, those metrics lean heavily on LLM judges. They are excellent for a nightly quality job that posts trends to Slack. They are poor bouncers for a pull request. Move RAGAS to your scheduled analysis pipeline, not your merge blockers.
Arize Phoenix carries the Elastic License 2.0 and sits at a different intersection entirely. It connects distributed tracing with evaluation, giving you observability into why a model behaved a certain way. You want this when you are debugging a production incident or tracing a hallucination back to a bad retrieval chunk. You do not want a tracing tool deciding whether a junior developer’s feature branch can ship. Its architecture is built for insight, not binary gates.
MLflow Evaluate (Apache 2.0) inherits its pedigree from experiment tracking. It is heavy. Pulling it into a lean CI image adds startup time and dependencies that slow down every single job. If you absolutely must use it inside a pipeline, stick to its heuristic metrics for structural checks. Even then, you are fighting the framework’s fundamental design. MLflow wants to log runs and compare experiments across weeks. A merge queue wants a verdict in under a minute.
Practical Rules for Gating
If you take nothing else from this experiment, take these three rules.
First, gate structure, not vibe. You can enforce that an output is valid JSON. You can enforce that it contains required keys. You can enforce that a classification label belongs to an allowed enum. These checks are fast, cheap, and deterministic. You cannot reliably enforce that a summary is "friendly" or that a rewrite is "creative." Those qualities belong in human review or periodic batch evaluation, not in automated gates.
Second, if a score moves on unchanged input, demote it immediately. Run your evaluation suite twice against the exact same artifact. If any metric flips from pass to fail, it has lost its right to block a merge. Promote it to an advisory dashboard where variance is expected and tolerable.
Third, respect the exit code. A pretty HTML report with a red banner does not stop a merge. A nonzero exit code does. Your evaluation tool must speak the native language of your CI platform. Standard out is for humans. Exit codes are for machines.
The Takeaway
We are still early in figuring out how to test LLM-powered applications. The temptation is to treat evaluation like a human grading rubric: nuanced, contextual, and slightly subjective. That works in a research paper. It collapses in a merge queue.
After eight months of production traffic, my pipeline now runs Promptfoo for structural and schema assertions across services, and DeepEval for Python-side behavioral checks that map cleanly to pass-fail conditions. Everything else reports to nightly dashboards. The queue is stable. The signal is clean. The team trusts a red build again.
You do not need more metrics at your gate. You need fewer metrics that tell the truth every single time.
Based on original testing and write-up shared on Dev.to. For more discussions on building reliable AI systems, join the GyaanSetu community on Telegram.
