একজন স্থপতি ভবন বানান — দ্রুত, সুন্দর, বিশাল। কিন্তু কখনো মাপেন না। কত উঁচু? কত শক্ত? কত নিরাপদ? একদিন ভবন ধসে পড়ে। "আমি ভেবেছিলাম ভালো," তিনি বললেন। "ভাবলে হবে না," বিচারক বললেন। "মাপতে হবে।" যে মাপে না, সে অন্ধ। যে মাপে, সে দেখে।
An architect builds — fast, beautiful, massive. But never measures. How tall? How strong? How safe? One day the building collapses. "I thought it was good," he said. "Thinking isn't enough," the judge said. "Must measure." One who doesn't measure, is blind. One who measures, sees.
এই বই তোমাকে শেখাবে — LLM কীভাবে মাপতে হয়। কোন metric কখন? Accuracy, faithfulness, relevance, coherence। Human eval, LLM-as-judge, automated testing। Benchmark design, eval pipeline, regression detection। যে মাপে, সে উন্নত করে। যে মাপে না, সে অন্ধের মতো চলে।
This book teaches — how to measure LLMs. Which metric when? Accuracy, faithfulness, relevance, coherence. Human eval, LLM-as-judge, automated testing. Benchmark design, eval pipeline, regression detection. One who measures, improves. One who doesn't, walks blindly.
দশটি মাপ, ক্রমানুসারে। প্রতিটি মূল্যায়নের এক একটি স্তর।
দশটি মাপ পেরিয়েছ। LLM evaluation-এর প্রতিটি স্তর আয়ত্ত করেছ।
Metrics, task-specific eval, LLM-as-judge, human eval, benchmark design, eval pipeline, regression, production eval, RAG eval, এবং complete eval architecture।
এখন তুমি জানো — মডেল কীভাবে মাপতে হয়।
যে মাপে, সে উন্নত করে।