Case Study Torch

⏳ 27% More Engaged Admins by Turning a 2-Person Bottleneck Into an AI Layer

Jump to: The Situation · My Role · What I Did · The Outcome · Hindsight

Rather than adding headcount to a bottlenecked insight-generation process, I redesigned the underlying data model so AI could deliver what a two-person behavioral science team could never scale to cover.

THE SITUATION

Torch—a virtual coaching B2B2C SaaS startup—displayed admin dashboards which held a valuable but inert asset: raw learner feedback from satisfaction surveys, buried in a format no customer admin could act on. Turning that feedback into something usable required an account manager to manually synthesize it, and to get a reliable read, that synthesis was reviewed by Torch's behavioral science team—a team of two people supporting all accounts. The insight pipeline for the entire customer base ran through two humans.

The stakes went beyond internal efficiency. Torch's growth strategy depended on renewals, and renewals depended on a credible ROI story that required admins to see, and return to, the platform. Admin engagement with the dashboard was one proxy for whether customers believed Torch was worth renewing. As long as insight generation stayed manual, that story could only be told to a handful of accounts at a time, and admins had little incentive to log back in between account manager touch points.

At the same time, Torch's leadership capacities—its proprietary set of measurable leadership attributes—existed as marketing language more than as a connected data layer. Nothing in the platform systematically tied learner feedback back to those capacities, so there was no way to benchmark progress or tell a consistent growth story across accounts. Fixing the bottleneck and activating the leadership-capacity data model were the same problem.

Torch admin dashboard, before
Torch admin dashboard, after

MY ROLE

I led this project end-to-end across design, engineering, account management, customer success, and Torch's behavioral science team. I owned the product direction including the decision to abandon two summarization approaches before finding the one that worked. And I was responsible for translating a two-person team's expertise into a system that could operate at the scale of Torch's full customer base.

WHAT I DID

Built a single source of truth for Torch's leadership capacities.

Before any AI summarization could be trustworthy, every piece of application data needed to map back to one authoritative capacity model instead of scattered, account-by-account interpretations. I treated this as a systems problem first: one place to define and update the capacities, referenced everywhere else, so the insight layer built on top of it wouldn't inherit inconsistency.

Rejected two summarization approaches before landing on the right abstraction level.

My first attempt summarized raw feedback directly against the leadership capacities, which oversimplified verbatim responses into generic, unhelpful statements. Summarizing feedback on its own, without that structure, produced too many fragmented categories to be actionable. Both failures told me the right unit of analysis sat in between; specific enough to be credible and broad enough to aggregate.

Let customer discovery set the taxonomy's grain, not internal preference.

Early discovery surfaced a specific finding: admins didn't just want to know what the problems were, they wanted to see how urgent or prevalent each one was, and urgency only reads as credible when it's measured. That finding is what pushed the design toward a living taxonomy of 100 initial problem areas, reviewed quarterly, that every piece of feedback could be bucketed into. It let account managers show not just an issue, but how many learners across how many responses were raising it. This was the basis for the AI RAG model (Haiku selected over Sonnet for cost optimization) and a governance team to ensure problem areas were properly represented.

Cut the recommendations feature after customer feedback exposed its limits.

The original plan pushed further than surfaced insights to offer admins prescriptive recommendations for how to respond. Customer feedback made clear we couldn't account for enough organizational context to make those recommendations feel relevant or actionable. I scoped it out of the MVP rather than ship something that would undermine trust in the insights themselves.

Traded historical caching over speed to market.

Early in the build, caching was raised as a way to preserve historical reference points for admins. I chose not to build it, prioritizing time to market over cataloging and accepting that insights would recalculate on each page load or new survey submission, meaning an admin's view could shift week to week. It was a deliberate de-risking move: get real usage and learning in front of customers faster, rather than delay launch to solve a problem we hadn't yet validated.

Sequenced the launch through account managers before customers.

I soft-launched to the account management team first, giving them time to build trust in AI-generated outputs before they represented that insight to customers. Only after account managers were comfortable (treating internal adoption as a gate—not an afterthought) did we message and campaign the release to the broader customer base.

THE OUTCOME

MetricResult
Admin engagementMAU % of total admins increased 27%
Reporting pages stickiness (WAU/MAU) increased 70%
Admin satisfactionAdmin NPS increased 92%
Efficiency (account management)Hundreds of hours per quarter of manual insight-synthesis work eliminated
Efficiency (behavioral science)2-person behavioral science team fully offloaded from manual review; bottleneck removed across all Torch accounts
Retention / renewalsStrategic accounts had positive renewal traction and account managers had a benchmarked, quantifiable data story to support renewal conversations
ScaleLeadership-capacity data model activated for every account setting the foundation for Torch Spark
Torch admin insights results screenshot

HINDSIGHT

I would have reverse-engineered the account management and behavioral science team's existing manual insights far more rigorously before writing a single prompt. We did light discovery on what "good" insights looked like, but the real pattern—the specific way experts weighted and phrased urgency—only became clear after several rounds of building, testing with real feedback, and adjusting. Front-loading that reverse engineering wouldn't have eliminated iteration, but it would have compressed it, since we were effectively rediscovering expert judgment we already had access to.

I would have also revisited the caching decision now that the feature has proven out. Choosing not to cache was a deliberate trade favoring speed to market and faster real-world learning over building a historical reference layer, and it was the right call to validate the concept quickly. But the mechanism it created wasn't fully accounted for: because insights recalculated on every page load and every new survey submission, an admin opening the dashboard one week could see a different top-five list than the week before, with no explanation for the change. That was an acceptable cost while the feature was still proving itself; now that it's core to the renewal story, I'd add caching with visible versioning, so admins can see what changed and why instead of experiencing quiet drift.▫