Tag: workforce

  • Human or AI? What building an AI solutions catalogue taught me about pro’s and con’s of AI technology

    The trade-off between hand-checked data quality and AI-driven automation and why the right line between the two moves with scale.

    The setup

    For the past few months I’ve been building the AI Solutions Hub, a catalogue that documents which artificial intelligence insurers actually deploy in the Swiss market. The platform’s promise isn’t completeness at any cost; it’s reliability. Every entry rests on a source, and every source is graded for how much it can be trusted.

    The architecture behind it is quickly told: sources (insurer newsrooms, trade press, targeted web searches, manual tips) flow into an automation layer (scheduled crawler jobs and an interactive agent), pass through several AI quality gates, land in a Postgres database, and only then go public. Sounds linear. But the interesting part, and the real lesson isn’t in the architecture. It’s in a single seam: where does the machine stop and the human begin?

    The one sentence that governs everything

    A catalogue is worth only as much as the reliability of its weakest entry. Two naive positions follow from that sentence, and both are wrong:

    1. “A human has to check every entry.” – Highest quality, but doesn’t scale.
    2. “Just let the AI do it.” – Scales, but produces neatly formatted garbage.

    The whole craft lives in between. And here’s what I had underestimated: the right dividing line isn’t fixed. It shifts with volume.

    Why handwork doesn’t scale

    With five entries I read every source carefully, check every link, weigh every phrasing. That’s craft, and it’s good. With fifty it becomes a chore. With five hundred I lose the overview, not through carelessness, but because human attention is a finite resource. The crawler surfaces dozens of signals a week; nobody reads forty sources at 11 p.m. with the same care as at 9 in the morning.

    The subtler failure isn’t speed, it’s coverage: a duplicate resubmitted under a different name will almost inevitably slip past me in a list of two hundred entries. The human is the bottleneck and, worse, an unreliable bottleneck whose error rate climbs with fatigue.

    Why pure automation isn’t enough either

    The counter-test is sobering. Language models are confident and, at times, completely wrong. A concrete example from the project: I had executive-summary PDFs a chatbot had produced cleanly formatted, complete with citations. On inspection, roughly half the cited URLs were simply invented. Plausible, formatted, non-existent. Vendor marketing reads the same way as a real deployment: “AI-powered claims handling” except the text names no insurer that actually uses it. A model left to its own devices publishes exactly that: claims that look like evidence.

    The bridge: frameworks and thresholds

    Here’s the actual methodological insight. To automate a judgement, you first have to make it explicit.

    The Admiralty Code, a two-axis rating system from the intelligence world, in use for the better part of a century and today codified in NATO doctrine does exactly that. It breaks the diffuse question “how much do I trust this?” into two separate, nameable axes: the reliability of the source (A–F) and the credibility of the specific claim (1–6). An insurer press release picked up by independent trade press is a B1. A slick but unattributed landing page is an E4. Once the judgement is structured this way, a machine can propose it — it no longer has to sense it.

    The second lever is thresholds. A gut feeling (“this is probably relevant”) can’t be automated; a number can. On the platform, the relevance filter discards anything below a score of 0.75, and that value isn’t guessed, it’s derived from the data: every entry I ever approved scored ≥ 0.8; the junk I clicked away averaged 0.70. The threshold turns my past decisions into something operational.

    And third, the feedback loop: every rejection reason I note when discarding an entry flows back into the filter prompts as a negative example. The system gets a little stricter with each human decision. That’s how you raise the automation ceiling without lowering the quality floor.

    When the human is non-negotiable

    The rule of thumb that crystallised: automate the routine, escalate the exceptional. What stays with the human:

    • Defining the criteria themselves. That an entry only counts when a named insurer actively deploys the solution is a human stipulation, no model derives that on its own.
    • The top confidence tier. No algorithm awards itself a “verified.” A pipeline may propose a grade; the highest tier is a human-only decision.
    • The irreversible step. Publishing makes an entry public. I don’t hand that off without a safety net even at high AI confidence it only runs behind a multi-stage control gate.
    • Genuine edge cases with conflicting sources and the cases where the whole framing is wrong.

    The gate doesn’t decide the hard cases. It sorts: the clear ones pass through, the doubtful ones land on my desk.

    The counterintuitive part: when the model has the better overview

    And now the twist I didn’t see coming. “Human = quality, AI = speed” is too simple. For certain tasks, past a certain volume, the human is no longer the better guarantor of quality.

    The best example from the project: a duplicate that would have slipped through. Two entries described the same insurer deployment, one called “Claims Voicebot,” the other “AI voicebot for claims reports.” A fuzzy text match rated the name similarity at 0.33, far too low to trigger. And a human skimming a list of two hundred entries wouldn’t have connected the two either, the names are too different. What caught it was a semantic AI judgement that checked the new entry against every existing entry for the same insurer and concluded: this is the same solution.

    The point: at volume, consistency and total recall beat human attention. The model holds all five hundred entries in view at once, every time, without fatigue, without an off day. The human’s strength is judgement on the ambiguous single case; the model’s strength is the overview across the mass.

    The division of labour

    Better done by a humanBetter done by an LLM
    Defining the criteria (“what qualifies at all?”)Reading forty sources with steady care — at 11 p.m. as at 9 a.m.
    The irreversible sign-off (publishing publicly)Extracting structured fields and filling them in two languages
    Weighing conflicting edge casesApplying the same threshold identically — no drift, no off day
    Spotting when the whole framing is wrongCatching near-duplicates across languages (Claims ≈ claims reports)
    Awarding the top “verified” tierCross-reading every claim against the live source — on every entry

    The learning nugget

    The reflex to frame this as “human or machine” leads you astray. The real work is designing the seam between them. Three moves have proven their worth:

    1. Make the quality judgement explicit – with a framework like the Admiralty Code that turns “I trust this” into two nameable axes.
    2. Translate it into a threshold a machine can pass or fail.
    3. Reserve the human for the rare – definitions, edge cases, the irreversible.

    And the punchline that ties it together: the right dividing line shifts with volume. What needs a human at ten entries belongs automated at five hundred – and the human moves up a level: from checking individual rows to designing the system that checks them. Data quality doesn’t scale by checking more. It scales by engineering the checking itself well.


  • Beyond AI Use Cases: Why I created an AI Solution Hub

    Over the past two years, I have had countless discussions with executives, business leaders, and AI practitioners about the future impact of artificial intelligence as part of my role for AI Strategy, Portfolio & Steering.

    One question appears in almost every conversation:

    Which jobs will AI replace?

    While understandable, I believe this is increasingly the wrong question. A more useful question is:

    Which activities within a role will no longer require human effort
    because AI can perform them more effectively, consistently, and at scale?

    This shift in perspective fundamentally changes how organizations should think about AI transformation.

    AI is changing tasks before it changes jobs

    Recent research from leading institutions such as MIT, Harvard Business School, and McKinsey points in a similar direction.

    AI is not primarily replacing professions. It is progressively taking over specific tasks within professions.

    Jensen Huang, CEO of NVIDIA, recently described this distinction as the difference between a person’s tasks and their purpose.

    In many cases, AI can automate significant portions of information-intensive work while leaving the actual business responsibility with the human expert.

    Consider an insurance underwriter. The purpose of the underwriter is not reading documents. The purpose is making sound risk decisions.

    Yet a large portion of the role today still involves gathering information, reviewing reports, validating data, and preparing analyses.

    These are precisely the activities that AI is becoming increasingly capable of handling.

    The same applies to claims management, compliance, finance, legal, HR, and many other business functions.

    The future is not AI replacing humans

    In my view, the future is better described as a redistribution of work.

    Historically, knowledge workers spent significant time on:

    • Searching
    • Reading
    • Summarizing
    • Documenting
    • Reporting

    Today, AI is rapidly taking over these activities. As a result, human work shifts towards:

    • Judgment
    • Prioritization
    • Governance
    • Stakeholder management
    • Decision accountability

    This is particularly relevant in highly regulated industries such as insurance, where accountability and human oversight remain essential.

    The question therefore becomes:

    How do we systematically understand which capabilities AI can already perform, which capabilities are emerging, and where humans will continue to play a critical role?

    Understanding the evolution of AI capabilities

    When viewed from a strategic perspective, AI capabilities are evolving through several distinct stages.

    Predictive AI

    The first wave focused on prediction.

    Examples include:

    • Fraud detection
    • Customer churn prediction
    • Risk scoring
    • Pricing optimization

    Generative AI

    The second wave focused on content creation.

    Examples include:

    • Text generation
    • Document summarization
    • Translation
    • Image generation

    Reasoning AI

    We are now entering a phase where AI increasingly performs structured analysis and problem solving.

    Examples include:

    • Complex case assessment
    • Risk analysis
    • Compliance reviews
    • Decision support

    Agentic AI

    The next wave goes beyond analysis.

    AI agents are beginning to execute complete workflows.

    This includes:

    • Gathering information
    • Using software tools
    • Performing actions
    • Coordinating multiple systems
    • Escalating exceptions

    This is where AI starts moving from being an assistant towards becoming a digital workforce.

    The capability question becomes a leadership question

    The most important challenge is no longer technological.

    It is managerial.

    Leaders need to understand:

    • Which activities create value?
    • Which activities can be delegated to AI?
    • Which decisions require human accountability?
    • Which new skills become critical?

    Organizations that answer these questions effectively will likely outperform those that focus solely on technology adoption.

    Looking ahead

    I believe we are only at the beginning of a much larger transformation.

    The conversation will gradually move away from chatbots and isolated use cases.

    Instead, organizations will increasingly focus on orchestrating collaboration between humans and AI systems.

    Understanding this shift requires more than experimenting with new tools.

    It requires a structured understanding of AI capabilities, business value, governance, and organizational readiness.

    This is exactly why I have created the AI Solution Hub.

    It is already being used to provide a structured view of AI capabilities, business use cases, opportunities, limitations, and governance considerations across different domains.

    Why I created an AI Solution Hub

    The purpose is simple.

    Organizations are currently overwhelmed by thousands of AI products, copilots, agents, platforms, and use cases.

    At the same time, expectations often exceed reality.

    Some believe AI can already solve almost everything. Others underestimate how quickly capabilities are evolving.

    The AI Solution Hub aims to create transparency.

    It provides a structured overview of:

    • Existing AI capabilities
    • Emerging AI capabilities
    • Relevant business use cases
    • Opportunities
    • Limitations
    • Governance requirements
    • Risk considerations

    Most importantly, it helps separate hype from practical business value.

    Its purpose is not to track technology for its own sake.

    Its purpose is to help organizations better understand what AI can realistically do today, what is emerging, and where human expertise will remain indispensable. A key principle behind the platform is trust. As I discussed in my previous blog post, trustworthy AI decisions require trustworthy data. Therefore, the information and assessments within the hub are evaluated using a structured methodology inspired by NASA’s Technology Readiness Level (TRL) framework, helping organizations understand not only what is technically possible but also how mature and reliable a capability is in practice. In addition, the platform provides a market perspective by continuously monitoring AI solutions, vendors, and emerging trends, enabling leaders to make informed decisions based on both capability maturity and market developments.

    If you are exploring how AI can create value in your organization and want a more structured way to navigate the rapidly evolving AI landscape, I invite you to take a closer look at the AI Solution Hub and see how it can support your AI journey.

    www.ai.marketeq.net