Your model is trained. But is it actually good? Most ML engineers default to accuracy. Then wonder why their model fails in production. Here are 20 evaluation metrics — and when to actually use each one: Classification: - Accuracy → Balanced datasets only. - Precision → When false positives are costly. - Recall → When false negatives matter more. - F1 Score → Imbalanced datasets. Balances both. - ROC-AUC → Binary classification evaluation. - Log Loss → Probabilistic models. Penalizes confident wrong predictions. - Confusion Matrix → Error analysis. See exactly where it breaks. - Specificity → When detecting negatives correctly matters. - Balanced Accuracy → Uneven datasets. Don't trust plain accuracy here. Regression: - MAE → Simple, interpretable error measurement. - MSE → Penalizes larger errors more heavily. - RMSE → Error in original scale. Most interpretable. - R² Score → How much variance your model explains. - Adjusted R² → Feature-heavy models. Adjusts for complexity. - MAPE → Business forecasting. Error as a percentage. - Explained Variance → Model consistency evaluation. Clustering: - Silhouette Score → Cluster cohesion and separation. Cluster validation. - Davies-Bouldin Index → Lower is better clustering. NLP: - BLEU Score → Machine translation quality. - ROUGE Score → Text summarization quality. Accuracy is not a strategy. Picking the right metric for the right problem is. A model that looks great on accuracy can destroy real-world outcomes when the wrong metric guided its evaluation. Save this. 📌 Which metric do most engineers misuse? 👇
Strategies For Change Implementation
Explore top LinkedIn content from expert professionals.
-
-
Most teams pick metrics that sound smart… But under the hood, they’re just noisy, slow, misleading, or biased. But today, I'm giving you a framework to avoid that trap. It’s called STEDII and it’s how to choose metrics you can actually trust: — ONE: S — Sensitivity Your metric should be able to detect small but meaningful changes Most good features don’t move numbers by 50%. They move them by 2–5%. If your metric can’t pick up those subtle shifts , you’ll miss real wins. Rule of thumb: - Basic metrics detect 10% changes - Good ones detect 5% - Great ones? 2% The better your metric, the smaller the lift it can detect. But that also means needing more users and better experimental design. — TWO: T — Trustworthiness Ever launch a clearly better feature… but the metric goes down? Happens all the time. Users find what they need faster → Time on site drops Checkout becomes smoother → Session length declines A good metric should reflect actual product value, not just surface-level activity. If metrics move in the opposite direction of user experience, they’re not trustworthy. — THREE: E — Efficiency In experimentation, speed of learning = speed of shipping. Some metrics take months to show signal (LTV, retention curves). Others like Day 2 retention or funnel completion give you insight within days. If your team is waiting weeks to know whether something worked, you're already behind. Use CUPED or proxy metrics to speed up testing windows without sacrificing signal. — FOUR: D — Debuggability A number that moves is nice. A number you can explain why something worked? That’s gold. Break down conversion into funnel steps. Segment by user type, device, geography. A 5% drop means nothing if you don’t know whether it’s: → A mobile bug → A pricing issue → Or just one country behaving differently Debuggability turns your metrics into actual insight. — FIVE: I — Interpretability Your whole team should know what your metric means... And what to do when it changes. If your metric looks like this: Engagement Score = (0.3×PageViews + 0.2×Clicks - 0.1×Bounces + 0.25×ReturnRate)^0.5 You’re not driving action. You’re driving confusion. Keep it simple: Conversion drops → Check checkout flow Bounce rate spikes → Review messaging or speed Retention dips → Fix the week-one experience — SIX: I — Inclusivity Averages lie. Segments tell the truth. A metric that’s “up 5%” could still be hiding this: → Power users: +30% → New users (60% of base): -5% → Mobile users: -10% Look for Simpson’s Paradox. Make sure your “win” isn’t actually a loss for the majority. — To learn all the details, check out my deep dive with Ronny Kohavi, the legend himself: https://lnkd.in/eDWT5bDN
-
Everyone’s excited to launch AI agents. Almost no one knows how to measure if they’re actually working. Over the last year, we’ve seen brands launch everything from GenAI assistants to support bots to creative copilots but the post-launch metrics often look like this: • Number of chats • Average latency • Session duration • Daily active users Useful? Yes. But sufficient? Not even close. At ALTRD, we’ve worked on AI agents for enterprises and if there’s one lesson it’s this: Speed and usage mean nothing if the agent isn’t solving the actual problem. The real performance indicators are far more nuanced. Here’s what we’ve learned to track instead: 🔹 Task Completion Rate — Can the AI go beyond answering a question and actually complete a workflow? 🔹 User Trust — Do people come back? Do they feel confident relying on the agent again? 🔹 Conversation Depth — Is the agent handling complex, multi-turn exchanges with consistency? 🔹 Context Retention — Can it remember prior interactions and respond accordingly? 🔹 Cost per Successful Interaction — Not just cost per query, but cost per outcome. Massive difference. One of our clients initially celebrated their bot’s 1 million+ sessions - until we uncovered that less than 8% of users actually got what they came for. That 8% wasn’t a usage issue. It was a design and evaluation issue. They had optimized for traffic. Not trust. Not success. Not satisfaction. So we rebuilt the evaluation framework - adding feedback loops, success markers, and goal-completion metrics. The results? CSAT up by 34% Drop-off down by 40% Same infra cost, 3x more value delivered The takeaway: Don’t just measure what’s easy. Measure what matters. AI agents aren’t just tools - they’re touchpoints. They represent your brand, shape user experience, and influence business outcomes. P.S. What’s one underrated metric you’ve used to evaluate AI performance? Curious to learn what others are tracking.
-
From Ceremony Manager to Value Driver: My Scrum Master Transformation 📊 I'll never forget my first retrospective as a new Scrum Master – the team was going through the motions, but we had no data to back up our feelings. That moment sparked my journey from simply facilitating ceremonies to truly driving team improvement through meaningful metrics. These KPIs transformed how I coach teams and demonstrate value to stakeholders. Here are the metrics that changed everything: 📈 Sprint Velocity & Burndown: Tracking consistency and predictability in delivery. ⏱ Cycle Time & Lead Time: Measuring flow efficiency from request to delivery. 🐛 Defect Density & Escaped Defects: Focusing on quality and technical excellence. 😊 Team Happiness & Collaboration: Because culture eats strategy for breakfast. 🎯 Sprint Goal Success Rate: Ensuring we're delivering valuable outcomes, not just outputs. These metrics don't just measure performance – they create conversations that drive continuous improvement.
-
Everyone says “change is happening” But how do you know it’s actually working? Change initiatives are easy to start. Harder to measure. Without clear indicators, leaders guess if progress is real And guesswork rarely works Top change leaders track these metrics to stay ahead: 1/ Achievement → How close did we get to our change goals → Focus on learning first, then performance Example: % of project milestones met vs. planned 2/ Completion → How well did we execute on schedule, scope, and budget → Example: Tasks finished on time and within budget 3/ Acceptability → Stakeholder satisfaction with the process and solution → Example: Survey scores, qualitative feedback 4/ Engagement → How involved are teams and stakeholders in the change → Example: Attendance in workshops, participation in feedback sessions 5/ Adoption → Are people actually using new systems, behaviors, or processes → Example: % of employees actively using a new tool or workflow 6/ Sustainability → Are changes sticking over time or fading → Example: Reassess behaviors 3–6 months post-change 7/ Impact → The measurable difference on business outcomes → Example: Efficiency gains, revenue growth, or error reduction Stop hoping for progress. Start proving it. P.S. Which of these metrics do you track most closely in your change initiatives? -- Follow me, Daniel Lock, for practical tips for leading change, consulting & thought leadership
-
𝗪𝗵𝘆 𝗧𝗿𝗮𝗱𝗶𝘁𝗶𝗼𝗻𝗮𝗹 𝗜𝗥 𝗠𝗲𝘁𝗿𝗶𝗰𝘀 𝗙𝗮𝗶𝗹 𝗶𝗻 𝘁𝗵𝗲 𝗔𝗴𝗲 𝗼𝗳 𝗥𝗔𝗚 𝗦𝘆𝘀𝘁𝗲𝗺𝘀 We've been evaluating information retrieval systems in the same way for decades, using metrics such as nDCG, MAP, and MRR, which were designed with human users in mind. But there's a fundamental problem: LLMs are not humans. In our new paper, "𝘙𝘦𝘥𝘦𝘧𝘪𝘯𝘪𝘯𝘨 𝘙𝘦𝘵𝘳𝘪𝘦𝘷𝘢𝘭 𝘌𝘷𝘢𝘭𝘶𝘢𝘵𝘪𝘰𝘯 𝘪𝘯 𝘵𝘩𝘦 𝘌𝘳𝘢 𝘰𝘧 𝘓𝘓𝘔𝘴," we identify two critical misalignments between traditional metrics and RAG reality: 1. 𝗛𝘂𝗺𝗮𝗻 𝘃𝘀. 𝗠𝗮𝗰𝗵𝗶𝗻𝗲 𝗣𝗼𝘀𝗶𝘁𝗶𝗼𝗻 𝗗𝗶𝘀𝗰𝗼𝘂𝗻𝘁. Traditional metrics assume users scan results sequentially, with decreasing attention down the list. But LLMs process all retrieved documents holistically and exhibit unique positional biases (like the "lost-in-the-middle" effect) that don't follow human patterns. 2. 𝗛𝘂𝗺𝗮𝗻 𝗥𝗲𝗹𝗲𝘃𝗮𝗻𝗰𝗲 𝘃𝘀. 𝗠𝗮𝗰𝗵𝗶𝗻𝗲 𝗨𝘁𝗶𝗹𝗶𝘁𝘆. Classic metrics treat all irrelevant documents equally—as neutral noise. However, in RAG systems, irrelevant documents aren't only useless but also actively harmful. Some passages distract LLMs, triggering hallucinations and degrading answer quality, while others have minimal impact. 𝗢𝘂𝗿 𝗦𝗼𝗹𝘂𝘁𝗶𝗼𝗻: UDCG (Utility and Distraction-aware Cumulative Gain) We introduce a new metric designed explicitly for RAG evaluation that: • Captures the positive contribution of relevant passages • Quantifies the negative impact of distracting passages • Uses an LLM-oriented positional discount • Directly optimizes correlation with end-to-end answer accuracy Key Results: ✅ Up to 36% improvement in correlation with RAG accuracy compared to traditional metrics ✅ Validated across five datasets (NQ, TriviaQA, PopQA, BioASQ, NoMIRACL) ✅ Tested on six different LLMs (Llama 3B-70B, Mistral, Qwen, Gemma) ✅ Robust across different context lengths and model families 𝗪𝗵𝘆 𝗧𝗵𝗶𝘀 𝗠𝗮𝘁𝘁𝗲𝗿𝘀. If you're building or optimizing RAG systems, you might be using the wrong evaluation metrics. Optimizing for nDCG or MAP doesn't necessarily improve what matters most: your LLM's ability to generate correct answers. UDCG provides a more reliable signal for evaluating retrieval quality in LLM-powered systems, helping practitioners build better RAG applications. 𝗧𝗵𝗲 𝗕𝗼𝘁𝘁𝗼𝗺 𝗟𝗶𝗻𝗲: As AI systems evolve from serving human consumers to powering machine consumers, our evaluation methodologies must evolve too. This work is a step toward aligning IR evaluation with the realities of modern LLM-based applications. • 📄 Read the full paper: https://lnkd.in/dCBYA_mw • 💻 Code and data available: https://lnkd.in/dB9BKDb7 What are your experiences with RAG evaluation? I'd love to hear your thoughts and challenges in the comments. #RAG #InformationRetrieval #LLM #AI #NLP A special thank you goes to my amazing co-authors: Giovanni Trappolini, Florin Cuconasu, Simone Filice, and Yoelle Maarek
-
Moving From Metrics to Meaningful Action Tracking metrics like employee satisfaction or turnover is important—they show where you stand. But numbers alone won’t create change. Metrics tell you what’s happening, but they don’t explain why or what to do next. Pairing KPIs (Key Performance Indicators) with OKRs (Objectives and Key Results) can help. KPIs give you the data, while OKRs tie that data to actionable goals that reflect impact—and asking why ties it altogether. For example, instead of just tracking eNPS (Employee Net Promoter Score), ask: • Why are we measuring this? • What does it tell us about our workplace? • How can we use it to create meaningful improvements for employees? By focusing on the connection between metrics, meaning and action, you can move from simply evaluating performance to transforming your organization. How are you linking what you measure to what you achieve—can your team tell you the impact and why we are tracking specific metrics? #KPIs #OKRs #EmployeeExperience #WorkplaceImprovement
-
"You can’t manage what you don’t measure." Yet, when it comes to change management, most leaders focus on what was implemented rather than what actually changed. Early in my career, I rolled out a company-wide process improvement initiative. On paper, everything looked great - we met deadlines, trained employees, and ticked every box. But six months later, nothing had actually changed. The old ways crept back, employees reverted to previous habits, and leadership questioned why results didn’t match expectations. The problem? We measured completion, not adoption. 𝗖𝗼𝗻𝗰𝗲𝗿𝗻: Many organizations struggle to gauge whether change efforts truly make an impact because they rely on surface-level indicators: → Completion rates instead of adoption rates → Project timelines instead of performance improvements → Implementation checklists instead of employee sentiment This approach creates a dangerous illusion of progress while real behaviors remain unchanged. 𝗖𝗮𝘂𝘀𝗲: Why does this happen? Because leaders focus on execution instead of outcomes. Common pitfalls include: → Lack of accountability – No one tracks whether new processes are being followed. → Insufficient feedback loops – Employees don’t have a voice in measuring what works. → Over-reliance on compliance – Just because something is mandatory doesn’t mean it’s effective. If we want real, measurable change, we need to rethink what success looks like. 𝗖𝗼𝘂𝗻𝘁𝗲𝗿𝗺𝗲𝗮𝘀𝘂𝗿𝗲: The solution? Focus on three key change management success metrics: → 𝗔𝗱𝗼𝗽𝘁𝗶𝗼𝗻 𝗥𝗮𝘁𝗲 – How many employees are actively using the new system or process? → 𝗣𝗲𝗿𝗳𝗼𝗿𝗺𝗮𝗻𝗰𝗲 𝗜𝗺𝗽𝗮𝗰𝘁 – How has efficiency, quality, or productivity changed? → 𝗨𝘀𝗲𝗿 𝗦𝗮𝘁𝗶𝘀𝗳𝗮𝗰𝘁𝗶𝗼𝗻 – Do employees feel the change has made their work easier or harder? By shifting from "Did we implement the change?" to "Is the change delivering results?", we turn short-term projects into long-term transformation. 𝗕𝗲𝗻𝗲𝗳𝗶𝘁𝘀: Organizations that measure change effectively see: → Higher engagement – Employees feel heard, leading to stronger buy-in. → Stronger accountability – Leaders track impact, not just completion. → Sustained improvement – Change becomes embedded in the culture, not just a temporary initiative. "Change isn’t a box to check—it’s a shift to sustain. Measure adoption, not just action, and you’ll see the impact last." How does your organization measure the success of change initiatives? If you’ve used adoption rate, performance impact, or user satisfaction, which one made the biggest difference for you? Wishing you a productive, insightful, and rewarding Tuesday! Chris Clevenger #ChangeManagement #Leadership #ContinuousImprovement #Innovation #Accountability
-
Lessons From the Frontlines: Evaluating RAG and LLMs for Real-World Results 1. 🎯 Anchor Everything in the Business Problem Teams often rush to optimize technical metrics, but the real question is: does this solve a business pain? Early in our journey, we celebrated high retrieval scores—until we realized customers still struggled to get clear answers. The lesson: define success with stakeholders before building. If the outcome doesn’t move a business metric, it’s the wrong target. 2. 🔄 Evaluation Is a Continuous Cycle The temptation is to treat evaluation as a milestone. In practice, it’s a loop. Each deployment reveals new user behaviors, new edge cases, and new opportunities for improvement. The most valuable insights come after launch, not before. The process: deploy, observe, learn, refine—repeat. 3. 🧩 Choose Metrics That Reflect Reality Technical metrics are necessary, but not sufficient. For RAG, the focus is on: Are we retrieving the right information? Are generated answers grounded in facts? Are users satisfied and returning? Is there measurable business impact? If a metric doesn’t tie to a business outcome, it’s deprioritized. 4. 🧪 Balance Offline and Online Evaluation Offline tests (precision, recall, latency) are fast and safe, but they don’t capture the full picture. Real-world use always surfaces surprises. One deployment looked flawless in the lab, but failed when users asked questions outside our test set. Only live data revealed the gaps. The approach: validate offline, but trust online feedback. 5. 🗣️ Communicate With Clarity and Context Data without context is noise. When sharing results, the focus is on the story: what changed, why it matters, and how it impacts the business. If a technical metric drops but user engagement rises, that’s a win. The narrative always ties back to business goals. 6. 🔁 Iterate Relentlessly The first version rarely gets it right. User feedback is the most valuable input—sometimes uncomfortable, always instructive. When users flagged irrelevant answers, the team didn’t defend the model. We adjusted retrieval, retrained, and improved. Each cycle brought us closer to the mark. Feedback is not a threat; it’s a guide. 7. 📝 Document the Journey Institutional memory is fragile. Every experiment, every decision, every lesson is logged. 8. Outcome-First Development (OFD) means starting every feature or model change with a clear plan for how success will be measured, not treating evaluation as an afterthought. This approach aligns engineering, product, and business teams, ensuring everyone builds with the end in mind. OFD helps catch issues early and keeps progress visible and actionable for all stakeholders. This discipline prevents repeated mistakes, accelerates onboarding, and ensures progress is cumulative, not circular. #RAG #LLM #AILeadership #AIEvaluation #BusinessImpact #MachineLearning #ProductStrategy #LessonsLearned
-
8 metrics run your maintenance program, yet most teams only track 3. Typically, the 3 chosen are those that make the monthly report look good: PM compliance, schedule compliance, and cost. However, each metric in isolation can be misleading. For instance, 95% PM compliance seems impressive until you examine MTBF and discover that the same repairs are being performed on the same equipment every 6 weeks. You're completing work orders, but not effectively fixing issues. Low maintenance costs may appear efficient until asset availability drops below 85%, resulting in production losses. This isn’t efficiency; it’s underinvestment dressed up for the budget meeting. Fast MTTR may indicate responsiveness, but it often means that the team is repeatedly fixing the same failures. The most effective maintenance teams I’ve encountered monitor 8 metrics together: - PM compliance - Schedule compliance - Work order backlog - Mean time to repair (MTTR) - Mean time between failures (MTBF) - Asset availability - Maintenance cost as % of RAV - Emergency work percentage These metrics are analyzed as a system rather than a scoreboard. Execution metrics (PM compliance, schedule compliance, backlog) indicate if the team is performing the work. Reliability metrics (MTTR, MTBF, availability) reveal if the work is effective. Financial metrics (cost, emergency %) assess if spending is appropriate or merely reactive. When PM compliance is high but MTBF remains flat, it signals that PMs require revision. When costs are low but emergency work exceeds 15%, it suggests maintenance is being deferred under the guise of savings. If availability appears satisfactory but backlog is increasing, it indicates future borrowing. Every metric has a partner that ensures accountability. Teams that excel in this area don’t possess superior data; they ask better questions. Keep this in mind for your next planning meeting. It’s crucial when someone presents a single KPI as progress. What combination of metrics provides the clearest insight into your program's health? I would genuinely like to hear what you track.