Beyond the Lab: Why Your Project Passed Testing, But Died in the Field

Share
Beyond the Lab: Why Your Project Passed Testing, But Died in the Field
You tested technology. But the field tests your assumptions.(Image: Perplexity AI, 2026)

*DestructionDesk | Issue #38 | 08.09.2026*

Last issue we ran the autopsy on the market's assumptions. This issue, we turn the scalpel inward. Twenty-five years spent as consultant, entrepreneur, coach, judge, and evaluator across deep tech, medtech, and corporate innovation programs — and the same failure signature keeps recurring: not bad technology, but an assumption nobody named until reality tested it, usually at the point of maximum sunk cost. None of what follows comes from a controlled study — it's what that many years of sitting on different sides of the same table looks like when written down honestly, flattering conclusions included.

1. Introduction

The Technology Readiness Level (TRL) methodology was developed by NASA in the 1970s to assess the maturity of aerospace technologies. It has since been adopted—and adapted—across a wide range of industries, including defense, energy, medtech, and increasingly digital and service sectors such as insurance and finance.

Despite its widespread use, our experience suggests that TRL does not map cleanly to these new contexts. The terms "Lab," "pilot," and "operational environment" mean very different things in a pharmaceutical trial, a food production line, a software app store, and an insurance claims system. What counts as "validated" in one setting may be entirely insufficient in another.

We attempted to apply TRL across these industries—and found ourselves repeatedly confronting the same gaps: assumptions left implicit, handovers that failed to capture critical context, component-level TRL conflated with system-level readiness, and governance structures that could not manage the portfolio of projects effectively.

This article is a report from that work. It is not a solution. It is a diagnosis—and a set of pointers for practitioners who are navigating the same challenges.

2. The Problem as Practitioners See It

Across industries, we observe a recurring pattern:

First, the language of TRL is applied inconsistently. A "Lab" in one organization means a working prototype on a bench; in another, it means a pilot with real users. A "pilot" in medtech may involve regulatory oversight and clinical endpoints; in digital, it may mean a soft launch on an app store. These differences are not merely semantic—they affect how resources are allocated, how success is measured, and when projects are considered "ready."

Second, the transition from Lab to pilot to production is often treated as a linear progression. In practice, it is rarely linear. Regulatory hurdles, user feedback, integration challenges, and shifting business priorities force projects to loop back, adjust, or stall. The TRL scale does not capture this well.

Third, and most critically, the assumptions underlying each stage are seldom made explicit. Teams test under controlled conditions, curate the environment to remove "easy" mistakes, and document what they did—but not what they assumed about the environment, the user, the regulator, or the integration context.

The result is a predictable pattern: the technology works in the Lab, passes the pilot, and then breaks in production—not because the technology is bad, but because an assumption that had quietly held up until then stopped holding, and it surfaced too late to absorb the cost. The failure isn't traceable to one clean root cause. What repeats is the timing: the assumption breaks after the money, the timeline, and the organizational credibility are already committed—not before.

3. A Framework for Understanding

What emerged from our cases is a way of framing the challenge. We found it useful to think about three interconnected dimensions:

3.1 Technical Readiness (TRL)

How mature is the core technology or component? This is the original NASA scale. A common pitfall is that TRL is often applied only to the innovation at hand—not to the full system. A component supplier may claim TRL 7, but the system into which it is being integrated may only be at TRL 4 or 5.

3.2 System Integration Readiness

How well does the technology integrate with its operational environment—including legacy systems, supply chains, and other components? A component may be technically mature, but system-level integration may reveal issues that lower the effective TRL.

3.3 Contextual Readiness

How well does the technology perform under real-world constraints—including regulatory frameworks, user behavior, and organizational processes? Regulatory approval and end-user validation often surface issues that were not visible in controlled testing. Moreover, the operational environment may require changes to the technology itself—not just validation of what was already built.

Depending on the configuration of the operational environment and the tasks to be performed simultaneously, the workflow may need to be altered, or the technology format or delivery mechanism may need to change. For example, a sensor reading designed for a second or third screen may need to be integrated into a head-up display to be useful in a real operational setting.

These dimensions interact. A technology may be technically mature but fail in system integration. It may integrate well but be misaligned with regulatory or user expectations. Or it may require reconfiguration to fit the operational environment—not because the technology is flawed, but because the environment demands it.

This is a framework in the sense of a conceptual lens—a way of organizing what we have observed. It is not a tested methodology. Nevertheless, it is a set of distinctions that we find useful for diagnosing where projects get stuck.

Cautionary Note: Workflow vs. Technology

A recurring tension across our cases is the relationship between technology and existing workflows. Two ERP rollouts a few years apart illustrate it well: a BPCS implementation between 1993 and 1998, and a SAP pilot in 1999. In both, organizations faced the same choice—adapt the workflow to the system, or adapt the system to the workflow. The decision was not always made on objective grounds. Users sometimes demanded that new systems replicate old processes—not because it was better, but because it was familiar.

The (apocryphal) Ford quote captures this dynamic: "If I had asked my customers what they wanted, they would have said faster horses." Ford never said it—but the line survives because the dynamic it describes is real. The point is not to ignore users. It is to listen critically—to distinguish between genuine requirements and habitual preferences.

Our framework does not resolve this tension. But it highlights it. Contextual Readiness is not just about whether the technology works in the environment. It is about whether the environment is willing—and able—to adapt to the technology, and vice versa.

4. What We Have Learned from Practice

A pattern that recurs across our cases is that critical information—assumptions about the test environment, the user, the regulatory context, or the integration landscape—is seldom made explicit. It is not that teams deliberately withhold it. It is that these assumptions function like unspoken rules: embedded in how the work was set up, but not captured in the documentation, not raised in reviews, and not probed during handovers.

The consequence is predictable. When the technology moves from a controlled setting into a live operational environment, these unspoken assumptions confront reality. What looked like a successful pilot turns out to have been successful only under a specific, curated set of conditions that no longer apply. Whether the underlying reason is a regulatory shift, a user habit, or an integration detail varies by case. What does not vary is when it surfaces: after the project is already committed, not while it could still be cheaply corrected.

This is not about negligence or bad practice. It is a structural feature of how innovation projects are typically managed. The information exists, but it is distributed across people, phases, and documents. It is never assembled in one place, and it is rarely tested explicitly before the next phase begins.

This observation is not drawn from a controlled study. It is a pattern seen across 25 years of work as consultant, entrepreneur, coach, judge, and evaluator—roles that each expose a different failure point in the same cycle. We offer it as practitioner pattern-recognition, not as measured data, and leave it to readers to test against their own cases.

5. Open Questions and Future Directions

There are still open questions. Our sample is drawn primarily from deeptech, medtech, and large corporate settings. It is not clear to what extent these patterns apply to digital-first industries, startups, or highly regulated sectors like insurance—though early signals suggest similarities.

We also do not yet have a systematic way to capture and test assumptions across the TRL progression. The AI-assisted portfolio scenario work is speculative at this stage and requires further exploration.

6. Conclusion

TRL is a useful crutch to break a complex reality into manageable pieces—so we can plan, budget, and review, which holds real value. But the map is not the terrain. No matter how carefully we define the stages, there will always be a gap between what we tested and what we assumed. Between the certificate on the wall and the number on the screen. Between the pilot that passed and the market that did not notice.

So instead of a set of recommendations, we leave this issue with the questions the pattern keeps raising for us:

  • Why do these assumptions surface at the point of maximum sunk cost, and never earlier?
  • If the information needed to catch the assumption already existed somewhere in the organization, what would it take to assemble it before the next phase—not after the failure?
  • Is closing loops between those who build, those who operate, and those who use, actually a documentation problem, or an incentive problem?
  • Who decides that a Lab is ready to become a pilot, and what do they lose—politically or financially—by asking harder questions before that decision?
  • Not every Lab should become a pilot, and not every pilot should become a product. Who in your organization has the authority to say so, and does anyone ever exercise it?
  • If technical readiness alone cannot justify the hard decisions—fund, kill, or scale—what would a portfolio review look like that treated unexamined assumptions as a risk category in its own right?

The details will differ by industry. The questions, we suspect, will not.


Destruction Desk
We perform autopsies on innovation’s failed assumptions.


This newsletter was edited by Manfred Lueth.


You received this email because you signed up for this newsletter from DestructionDesk.com.
To stop receiving this newsletter, unsubscribe or manage your email preferences