Skip to content
Log in
— Episode 13 · 38 min

Metrics to track or toss

Craig Stoss argues that most support data is not worth keeping, most migrations are not worth doing, and capacity planning is far simpler than teams make it. He walks through occupancy, shrinkage, and the handful of metrics that actually change business decisions.

June 10, 2025 · Jen Weaver with Craig Stoss, VP of Partner Solutions, Kodif

— Takeaways —

What you’ll learn from this episode

  • Rigid fields are the first mistake. Restricted single-selects add friction for agents and force square pegs into round holes.
  • Do not migrate your data. Pull out volume trends, basic categorization, and handle time, and throw out the rest — a ticket's priority from 12 months ago will not change any decision you make.
  • Capacity planning needs only two numbers — volume, and the time it takes to handle that volume. Everything past that is optional.
  • More variables means less accuracy, because every variable carries its own variance. Pick a method, keep it as simple as the business allows, and stick to it.
  • 100% occupancy burns people out and produces bad data. Start around 80–85%, then subtract internal shrinkage (meetings, breaks) and external shrinkage (leave).
  • Track what changes the business — the share of tickets that become bugs tells engineering how many people to hire. The definition of a bug does not drift the way categories do.
  • Before AI could do it, his team used six tag dimensions with a minimum of three per ticket. Fluid enough to evolve with the product, structured enough to report on.
— Chapters —
  1. 0:00Introduction
  2. 2:22The two biggest data mistakes
  3. 4:44Why data migrations waste your time
  4. 7:04How data gets unclean as you grow
  5. 9:25Don't force old data to look clean
  6. 11:45Where AI changes reporting
  7. 14:07What capacity planning actually needs
  8. 16:27The Erlang model, and not over-complicating it
  9. 18:47More variables means less accuracy
  10. 21:08Occupancy, and why 100% burns people out
  11. 23:28Internal and external shrinkage
  12. 25:49The law of averages
  13. 28:10Metrics that actually impact the business
  14. 30:30Six tag categories, minimum three per ticket
  15. 32:51What to keep when you migrate
  16. 35:12Throw out the email threads
  17. 37:32Recap
— Transcript —

Transcribed from the recording and edited for readability — false starts removed, product and proper names corrected.

Craig Stoss We focus a lot of time and effort on, well, how do I migrate my data between systems if that's required? I've outgrown my current help desk, I want to migrate to a more robust one — what data do I migrate?

And the answer, in my opinion, is: don't. It's not worth your time and effort. I would throw almost all of it out. I would say pull out some major trends that you need just to be able to support future decisions — that would include things like volume trending, maybe some basic categorization, or handle time numbers.

Jen Weaver Welcome back to Live Chat with Jen Weaver, where we dive into the exact moves of top support pros — what they're doing to simplify operations and elevate the customer experience.

Today we're talking with Craig Stoss, VP of partner solutions at Kodif. He's here to help you stop wasting time on bad data, rigid systems, and over-complicated capacity models. If you've ever felt stuck trying to make sense of messy historical data, or figure out how many people you really need on your team, this one's for you.

Before we get started — this is the very first time I'm ever showing this to the world, our QA tool, Supportman. So if you're listening to this podcast, head over to the YouTube link in the show notes to get a glimpse. Supportman sends real-time QA from Intercom to Slack with daily threads, weekly charts, and done-for-you AI-powered conversation evaluations. It makes it so much easier to QA Intercom conversations right where your team is already spending their day, in Slack.

All right, on to today's episode. I'm really happy to chat with you, Craig. Can you tell me a little bit more about yourself?

Craig Thanks for having me, Jen, really happy to be here.

My name is Craig Stoss, I am the VP of partner solutions at Kodif, and I have been in the customer experience industry for almost 25 years now. I've spent a lot of time in support. I have done everything from teach children computer camps to consult with Fortune 500 companies. I've led global teams for companies like Shopify and Arctic Wolf, and now I sit with Kodif, which is a CX agent AI for your support teams, to bring efficiency across all your different ticket types.

My background has always been, in my heart, the support side of the CX business — working with all those tools like Zendesk and chatbots and analytics tools, and things that every support team needs.

Jen That's fantastic. I'm so excited and honored that you're here with your experience. I think we could probably do 10 podcast episodes, but today I know you have a process about data that you want to talk about.

Data is something that we've gotten requests from our audience members a lot about, because it's tricky to deal with support data. So just to launch you into that discussion — what's the biggest mistake you see support teams making with their data?

Craig That's a good question. I think that there are two big mistakes.

One is making your data too rigid. By that I mean creating these drop-down fields, or fields that are restricted for certain circumstances, and really putting hard restrictions on how the data can be set and what data can be set at what times.

I think that's a huge mistake. It adds friction to the agent in order to use that data and to set it appropriately. And I think it also limits the usefulness of the data in the end, because you're trying to fit a lot of square pegs around holes, because the data doesn't allow you to use a round peg.

Jen Like single select fields?

Craig Single select fields is a great example. Those types of data I think are less useful, especially when it's restricted single select. If you can add new values dynamically, sure, maybe that's a better solution.

And then I think the second biggest mistake is a lot of support leaders, when they come in, or as companies grow, they want to migrate their data to a new system. Or they want to recreate the old scenarios inside some new tool, or some sort of revamp of an existing tool.

And you end up trying to spend all this time to figure out where all this data should sit, which you either are never going to look at again — or if you do look at it, now all of a sudden you're comparing apples to oranges.

I think there's so much time wasted thinking about data migration, whereas if you just really think about why this data exists, you don't need to do it. Those two mistakes make up for the biggest time sinks that support leadership sees when it comes to data.

Jen Thanks for being game for pop question time, I appreciate that. But I know you have this process you want to share with us about data. If I'm on a support team and I'm just beginning to think about data, what's my entry point?

Craig It's interesting. I think it's a stage thing.

At initial stages, the worry isn't the data. The worry is just, we need some tool to help us manage incoming tickets. Maybe we're a startup, maybe we're a small e-commerce mom and pop shop, whatever it is, but we need just something to manage the incoming volume. We don't care about reporting necessarily, we don't care about slicing our customers up by VIP versus non-VIP.

Jen We just don't want to use Gmail any more.

Craig Exactly, that's it. And so what happens is that because there's no vision — and I'm not even arguing that's a bad thing, not having that vision — you set your system up in a certain way, and then as you evolve, all of a sudden you start adding in new pieces of metadata.

If we look at a timeline, maybe at time zero plus five, all of a sudden you add this one field. And then at time five plus another 10, you add in some other metadata. And now all of a sudden you have this data but it's not consistent. It's not comparing apples to apples. Maybe that data from time five has evolved by time 20 so that it has new fields. And so now you end up with a set of data that is not complete, is not clean.

That's one path it goes. The other path is where someone wants to be smarter than the future, and they start to try to think about, well, how can I build that vision? So they want to over-engineer at the start — which again, I'm not faulting. But what happens is you're fast-paced, maybe you only have one engineer or one support agent that's helping you, they're not setting this data, they don't have time to fill that out. And it doesn't align with anything that actually happens in the future. Maybe your vision was wrong, for example.

So you end up almost always with unclean data. And I don't fault anyone for that, because I think that's how you have to start.

The process I would take is that when you are ready for data — when your business is at a point where you need to start using customer feedback, customer trending, you need to start doing capacity planning, you need to start building structured dashboards and reports so that you can report to the business what you're doing — the problem starts at this point, because in almost every scenario your data is going to be unclean. It's not going to be in a usable format.

The mistake that I see so often is that we try to force it to be clean. We either try to go back in history and change stuff, or we try to build assumptions into it, or we just report on it and tell the end user that's reading it, well, this isn't accurate, but this is where the trend is going. And that obviously has a lot of thumb-in-the-air type logic to it.

Jen I can feel the impulse to do that. If I'm just scrapping all my past data, I have nothing. I need that data — that's the impulse.

Craig The impulse is I need data. But if the data isn't representative of something useful, then I would argue you don't need data.

I've seen so many people who say, oh, there was a sudden spike in this type of ticket — and the root cause was literally that someone found out that that type of ticket was a selection and told the team, and they just started using it. It had nothing to do with the reality of the tickets.

Jen I've done that before. I've been a support specialist and just — I have to use this single select field, I have to do something, I'm in a hurry, my numbers matter to my compensation. So I just need to pick something and it's not representative. And I'm just embarrassed that I have been that problem.

Craig Every company, every leader feels this. Especially in support, where I'm sure everybody listening to this can resonate with your boss coming to you and saying, I need to know this exact number now. I need to know how many of our gold-star customers opened critical tickets over the last two quarters — give that to me by end of day today. And you're like, okay, I'll see. That's not a report that I have.

What's interesting is, and this is a bit off topic, but I think this stuff is getting easier now. Reporting tooling is getting a lot less static in the new AI world. There are tools that exist now that can go back and infer stuff, can make judgments.

And not only that — you don't even have to start building these reports based on fields as we're traditionally used to, in Salesforce and Zendesk and these tools, where you have to say, well, I want this field trended over this period of time, and I want it split based on this metric, and I want it filtered by only these types of customers. That's a very traditional way of thinking about reporting.

But there are now tools using AI where you can just describe what you want. Put in whatever your boss said to you — gold-star customers, last three quarters, that opened critical tickets — and it can go back and figure that out, and correlate data from multiple systems to make that intelligent report for you.

That's solving a lot of this problem, because you don't have to rely on that rigid data that I said earlier. Or Jen, to your point of, hey, just pick something because I think this is right. The AI is picking that something and being able to categorize that ticket in many different dimensions. So I think this problem is slowly going to go away.

Jen I can even see AI being able to look at the full text of a conversation and categorize it after the fact. That makes total sense to me.

From talking to support leaders, I'm getting the impression of what you're saying — which is we often come into organizations where we're maybe the first hire, or it's a new support team, and the data just hasn't been there. So what do we do from there? How do we land on a support team and then start from now forward?

Craig I think first and foremost the most important thing is to decide what data is likely to have some sort of impact to your business. What data is actually going to change something about your business?

I mentioned capacity planning. Capacity planning is almost entirely based on volume of tickets and the time it takes to handle those tickets. You don't need to know necessarily all the categories of tickets in order to get those two numbers out of a system. And that will impact your business. That's an important thing that you need to make sure is consistent and something that you can compare over time.

Jen So what are the data points for capacity planning that you recommend?

Craig It depends how robust you want to get, I suppose. But in general, literally you would measure number of incoming tickets or emails or phone calls, depending on how your omnichannel is set up. Chats is obviously huge right now. But you would measure that pure volume number from your system.

Maybe it is split by channel, in which case as you grow or expand different channels you can start to see which channels are trending better. So there might be one split in that, but I wouldn't split it much more than that. Capacity planning is really about the volume, and then the handle time for that.

Now, handle time is interesting because there are all sorts of ways to measure handle time. It's a metric that everyone measures, but everyone has to pick a way to measure it, and it's not consistent from company to company. It's not even consistent leader to leader within a company.

So you need to choose which way you think about it. Do you think about it as the actual time spent working a ticket? And does that include if the agent has to go and make some manual actions in another tool — go out to Shopify, or go out to an order management system and take action — does that time count? Does responding to that person count? If you're doing three chats at once, is that three separate handle times, or is that all just one unit of time? There's lots of stuff that goes into it.

For high volume places there's a methodology called the Erlang model. In fact, if you go to my website, stoss.ca, I actually have an Erlang calculator that I built to help leaders with some of these things, because I think it's an interesting way for high volume companies to do capacity planning. That requires a whole other set of data around when your tickets come in, what are your arrival patterns of tickets.

So to answer your question, it really depends on how you want to do capacity planning. The simplest is literally two metrics — volume, and the average time it takes you to handle that volume. It could be literally two numbers over a period of time, daily, weekly, monthly, quarterly, whatever it is.

But you can get much more advanced. You could split it by channel, you could split it by hour of the day, day of the week. It really depends on how in-depth you want to get. But the key is, for that specific use case, minimize that to something that's going to actually deliver impact to your business.

Jen I'm really interested in what you said about the calculator. You said there are many ways to calculate handle time and to deal with capacity planning, but it sounds like you have an opinion about what's the best way, because you built that calculator.

Craig I hesitate to use the word best.

At a previous company of mine we actually built three versions of the Erlang model, because we felt that there were three different levels your business had to be at to use different levels of the calculation.

The base one is what I said — trending volume and handle time, divided by number of people. There's your capacity.

And then the one that's on my website is literally, here is the exact number of people you need per hour, per day, in order to cover the volume that you're predicting. That gets pretty intense, and you'd probably have to be pretty high volume to start thinking in that direction. Because then if you have multiple channels, that adds another layer of complexity — it's like, well, you need five people to handle your email channel and you need 10 people to handle your phone channel, and how do you staff to make sure that that's minimized? Is it just five plus 10, or maybe the phone people can handle emails while they're not on the phone? Is it blended?

So my opinion is that you shouldn't over-complicate it. Going back to wasted time — and this might get me in trouble with some purists — but I think we spend so much time worrying about the calculation and making sure the numbers are accurate for something that we can't predict.

We're smart enough to predict within some amount of reasonable estimates. E-commerce is a great example of this. I work a lot with e-commerce companies at Kodif and we talk a lot about how many tickets are going to come from an order — what's the average volume of tickets based on my order volume? And therefore we know if order volume goes up for Black Friday, we know how many tickets are going to come, give or take.

But that's still very simple. I have people that try to extrapolate that on a weekly or daily basis, and they get so bogged down in these numbers that they lose accuracy because of the number of variables they put in. Because each variable has its own variance within it. So statistically, the more variables you use, and the more variance each of those variables has, the less accurate your last estimate is going to come out.

So my opinion is to figure out a method, make it as simple as your business requires — which again at high volume might mean a full Erlang hour-by-hour calculation, but certainly not every business needs that — but make it as simple as possible and then stick to it.

And if you discover a quarter or two in that there is some error in your calculation, then correct for it. But don't try to be that visionary of, I'm going to account for everything. I've had people try to account for, well, a person is sick one day a month and that typically happens on a Monday. No — what's the number of vacation and sick days that you see in a year? Spread that out over the year and that gives you your shrinkage number that you need to put into your calculations.

But it's an average. It'll never be exact. You might have 10 people sick one day and your business is literally crippled. There's no way to predict that. There's no methodology in the world that will predict that for you, at least not until AI gets a lot smarter.

So yeah, that's my opinion. Don't make it complex, because it will always be imperfect anyway.

Jen That makes a lot of sense. There's a piece of that that has this question burning in my brain.

I've talked to a lot of support leaders about metrics and capacity planning, and the pieces that you mentioned are crucial. But the other piece is, how much time out of a full-time individual contributor is ideal for them to be in the queue?

I've talked to leaders who say a percentage — for specialists, especially after they mature with the company and they've been there a couple of years, they need to be doing something else in addition to work with customers. And I've heard different ideal percentages of that time. Do you have thoughts about that, because that factors into that calculation?

Craig Oh, absolutely. And again, as with most answers in this world, the answer is it depends.

There are really two types of things that affect your full-time worker. Let's just assume an eight hour day, just so we have a starting point.

The first is occupancy. Occupancy is the amount of time that is spent in the queue directly. In some call centers it might be close to 100% — because you might be hanging up a phone and the next call just immediately rings into your ear, and then you hang that up. Think of the airlines and telecom companies. There's no downtime. When you are working, you are 100% working, you're never waiting for a ticket to come in.

And I would argue that that leads to burnout. I would argue that leads to unhappy employees. I would believe that leads to bad data, in the sense that you're obviously not going to log stuff as well. So that's probably not ideal.

I kind of use 80% as my starting point for most teams. Now, it can fluctuate. I ran a tier three team for a long time where the occupancy for new tickets was like 20%, because the tickets that they did get were so highly complex that they spent 80% of their time working on the existing tickets they had in their queue.

So it is going to depend on your team, but I tend to start with something around 80–85% as my goal. That way — and if you take that on an hourly basis — that's nine minutes of every hour that's kind of downtime, where they can wait for the next ticket, edit some notes, maybe follow up with another customer, go to the washroom, whatever it is. Just have some downtime.

And the second one is shrinkage. There are two types of shrinkage. There's shrinkage that's inherent to the business, usually called internal shrinkage. That's things like a team meeting. If you have a one hour team meeting per week, what percentage of time is spent there? If we assume a 40-hour work week, one hour is 1/40th of your time.

Maybe there are other meetings that someone attends — for example, you said as someone gets more senior maybe they attend an engineering meeting, or they attend a product meeting to understand what's happening in the product, or some sort of status update meeting, or maybe there's an all hands meeting that's once a month.

How I calculate this is I basically list out all the internal reasons why someone is not going to be working in the queue. That could be one-on-ones, all hands meetings, lunch breaks, break time — if you get two 15 minute breaks a day or whatever it might be. That is stuff that the business gives to you as part of just operating the business.

That number can vary. I see it around 10–15% somewhere in there of shrinkage of your time, but it could be lower depending on the nature of your business.

Then there's external shrinkage, which is the time that you take. That's like vacation, sick leave, that could encompass things like parental leave if you have a policy around that. It encompasses anything that you as an employee disrupt. And that number is usually a lower percentage overall, but obviously it has impact.

So if you take your shrinkage and you say, okay, my 40-hour work week is reduced by let's use 15% — so 40 hours is reduced to 34 hours of actual availability. And then my occupancy is 80%, so I reduce that number by a further 20%, which is another six and a bit, let's say seven. So that gets into 27 hours that someone is actually available in the queue on average.

Of course some weeks it might be closer to 36, some it might be lower than 27 because they take a vacation day. That's really the basics of that calculation.

And again, don't over-complicate it. Trying to figure out, oh, people take more vacation in the summer, so let's up it for the summer and lower it for the winter — why? What if I get sick in the winter? More people are sick in the winter than they are in the summer. Do you have to predict people's cold schedules and flu schedules?

It's just simple averages. And the law of averages is — especially if you have a big team, it's harder if you have a team smaller than 10 or 15 people — but if you have a 20, 30, 40, 100 person team, the law of averages says you're going to have the same number of people out every day sick, just because.

Jen They're human.

Craig They're human, it's just the nature of it.

And the ones that abuse that — I had an example where I had an employee who loved his football on Sundays, and loved it so much that he called in sick many days. That's a performance issue. That's not a problem with your capacity planning. That's an individual abusing a system that wasn't meant to accommodate for hangovers.

Jen In addition to capacity planning, what are some other metrics that maybe new support teams need to start thinking about moving forward — not going back to past data?

Craig So again, I'll bring it all back to what impacts the business.

We focus a lot on categorization of tickets, and maybe some levels of categorization need to always be there. For example, what percentage of tickets result in bugs? That's an interesting metric, especially in the SaaS world. In the product world, what percentage of tickets are related to a product arriving broken or damaged? What percentage of tickets are refund requests or return requests? Because those will have an impact on your business. There's a cost to those things.

Jen And like you were saying earlier, they're also manipulable. What percentage of tickets lead to bugs — we can put the engineering team on that and fix that over time, or at least diminish it.

Craig 100%. That's why I say start with what impacts your business. The priority of a ticket from two years ago is probably not going to impact your business. But if the number of bugs month over month has increased as your product complexity has gotten higher, that is an impact to your business. That's an impact to your customers, that's an impact to the number of engineers you need to hire.

Jen So support data is going to impact how many people engineering hires. That's something that people need to know and see that trend.

Craig And it's pretty cut and dry. That's something that typically remains apples to apples. The definition of a bug doesn't change. The category of where that bug is might change, because you might add new parts to your product, you might change parts of your product, you might pivot your entire company — so that data is probably less relevant over time. But the fact that you had a bug remains the same. And so as a percentage, that's a really relevant metric.

Before AI solutions to this problem, I had my team lay out six categories of tags — the ways that you could describe a ticket. They were things like the feature that was being used; the type of action, for example was it an edit action, a create action, a delete action, an update action; things like was it an end user, was it an admin, was it a different type of business user. We had six categories.

And the agreement with the team was that you would categorize a minimum of three categories per ticket that were relevant — because sometimes not all six were relevant, but you had to have a minimum of three.

That was excellent for data, because now you could start to see it. It was fluid, because the product evolved and we could start to add new categories as we learned about them, or as we saw these trends. It was a bit subjective, as is any human solution, but it worked quite well. We could report that data to other departments to say, look, our admins are struggling here, updates are failing at this particular feature, whatever it might be. We could be really descriptive.

AI can now do that. I won't say for free, but pretty easily, and certainly more accurately than a human could.

Jen And that saves specialist time and effort too, if they can just close a conversation and know that whatever tool you're using will gather that data.

Craig And then I think the last piece is, as your business expands, you might start a tier two team, or you might need to start understanding that specialization is important to you — so you have people who specialize in different areas of your product. If there's data you have on that to help you build that team in the future, that's useful. But again, if that's not available, don't force it.

Just pick a date and say, from this date forward, we know that we can categorize tickets in a way that says here's a specialization that we could potentially have.

There is some visionary stuff here that I'm not always a fan of. But at the same time, most support leaders have a gut feeling on, oh, this is a specialization, this is a completely different skill set. Maybe refunds and accounts is a completely different specialization than an app bug. Those are different skill sets, and eventually we likely will specialize to split those out, so that a refund is not handled by a technical bug analyst and a technical bug analyst isn't handling a refund.

So those are things that you would need to keep or start doing at a certain scale.

Jen That makes total sense. I know we're coming to the end of our time. Is there anything else you want to add to the data conversation to encourage support leaders? And also, if I'm a busy support leader and I need to throw out 90% of the many metrics, what should I focus my hour per week on?

Craig There are two things that I'd like to touch on.

One, I mentioned migration right at the top of the episode. We focus a lot of time and effort on, well, how do I migrate my data between systems if that's required? I've outgrown my current help desk, I want to migrate to a more robust one — what data do I migrate?

And the answer, in my opinion, is: don't. It's not worth your time and effort. You talk about throwing out 90% — I would throw almost all of it out. I would say pull out some major trends that you need just to be able to support future decisions. That would include things like volume trending, maybe some basic categorization or handle time numbers.

But you do not need to have the category of a ticket from 12 months ago, or the priority of a ticket from 12 months ago. It's not going to impact your business in any way. And even if it were, what you define 12 months in the future is not going to be the same, just because your business has evolved. Planning a migration at this scale, to me, is a waste of effort in 95%-plus of cases.

Now to answer your other question. If I were to throw out 90% of the data, I think I would start with everything that I know I won't ever look at again. And in most cases that's your email threads. You don't need to maintain an entire email thread from 12 months ago.

Now, you might want to — and this is what we're going to talk about in our next episode, AI and stuff like that — you might want to maintain summaries of that information, or word clouds of that information. But the actual email thread itself is irrelevant.

I would throw out almost any date-related data. We store things like deadline dates, or follow-up dates, or how many tickets lost SLA over a month. That's just not going to be relevant in the future, it's not going to change your business. That's a real-time performance metric. It's something that you use to coach an agent, or to inform a new hire, or worst case inform a layoff, to be honest. That stuff is not going to be useful outside of the real-time analysis.

Schedule information certainly can go. I've already talked about priorities.

So I would focus on the basics. If you want to keep some of these trends, use a business intelligence tool — Google Looker Studio or something to that effect — or even just an Excel spreadsheet. Put this trending data there so you can refer to it. But just be selective. And certainly the biggest chunk of it is all the back and forth email threads. No one needs to store that at any capacity.

Start fresh with new data, new systems that are able to categorize your tickets in a way that helps your business.

Jen That's refreshing. I remember being a support leader when we moved from one help desk to another, and I just felt like I need these threads for some reason. And I never used them. But I felt like — what if a customer comes back and says something and I have to refer to something? And it just never happened.

Craig Well, it won't happen. It doesn't happen.

Jen Thank you so much for being here and for chatting about data. I really appreciate it. Can't wait to have you on the podcast again.

Craig Thanks, Jen, I really appreciate it. Thank you for having me.

Jen Thanks so much to Craig for sharing his hard-earned insights. He also has a free Erlang calculator at stoss.ca — that's his website.

And as always, if you enjoyed today's episode, please hit that subscribe button, share it with a fellow support leader, and subscribe to our newsletter that's in the show notes to keep these tactical tips coming your way.

All right, so let's review the steps that Craig went over.

Step one, stop over-complicating your data structure.

Step two, don't waste time on historical data migrations.

Step three, choose a simple and consistent capacity model — 80% tends to make sense.

Step four, factor in real-world time for your team, not just ideal capacity.

Step five, only track data that drives meaningful business decisions.

Thanks for being here, and until next time, keep it simple and stay focused on what really moves the needle.

Under two minutes to live, no IT ticket required.

See pricing