Workshops & training

Get your team ready for the AI era

Most companies have already bought the licences. Almost none have changed how the work runs. We train executives, managers and practitioners on where AI actually holds up, then rebuild one real process around it inside the tools your team already uses.

13 modules · 32 contact hours · delivered on your own self-hosted instance

The idea the curriculum turns on

The capability boundary is jagged, and it is invisible

In a randomised experiment with 758 BCG consultants, the same model that sharply improved work on one side of its capability boundary degraded it on the other. Tasks that look equally difficult land on opposite sides.

That is why we do not teach prompts. Teams map the boundary against their own work, because it cannot be predicted from intuition and it moves every time a model ships.

Dell’Acqua et al., HBS Working Paper 24-013
Inside the frontier

25.1% faster

BCG consultants using GPT-4 completed 18 realistic consulting tasks 25.1% more quickly and produced results rated more than 40% higher in quality than a no-AI control group.

Outside the frontier

19 points worse

On a task deliberately placed outside the model’s capability boundary, the same consultants were 19 percentage points less likely to reach the correct answer than colleagues working without AI.

How we run it

Four commitments that make the training stick

Workflows, not tools

A tool demo teaches people to type into a box. We rebuild a process end to end with agents at the centre, because that is the change the research ties to actual earnings impact.

Your systems, your data

Every lab runs against your own Jira, GitHub, Linear or Notion on a self-hosted instance. Nobody leaves with a certificate and a sandbox they will never open again.

The boundary, taught honestly

We teach where AI degrades your work as carefully as where it improves it, and we teach when not to use it at all. Skeptical senior staff are the audience that decides whether this sticks.

Artifacts you keep

Workflow maps, eval sets, threat models, permission matrices, runbooks. Each lands in your repository under your licence, not in a vendor portal you rent.

Who it is for

Three tracks, because three audiences need different things

Selling the same deck to an executive and to an analyst is the most common way corporate AI training fails. The EU AI Act’s literacy obligation says the same thing in legal language: training has to account for the role and the context.

Half day + two checkpoints

Steering

C-suite, business unit heads, the executive sponsor who controls the budget line

Not a lighter version of the technical day. Leaders get portfolio triage, unit economics, regulatory exposure and decision rights — the things that decide whether a working pilot ever reaches production.

Leaves with

Three ranked candidate workflows with named owners, blockers and cost ceilings, plus a decision forum booked inside the real budget cycle.

26 contact hours

Workflow owner

Team leads, product and delivery managers, operations managers accountable for a process end to end

The core track. One real process, re-architected around agents, built, evaluated against a counterfactual, cost-ceilinged, threat-modelled and running against live events in shadow mode.

Leaves with

One workflow running, with a permission matrix, an operating runbook, a named operator and a written promotion path to production.

25 contact hours

Practitioner

Analysts, engineers, marketers, designers, customer operations and R&D staff

Depth on delegation, verification, context and evals — the skills that separate useful output from polished output nobody can trust. Governance and operations are covered as briefings, not builds.

Leaves with

Each participant ships one workflow they use weekly, plus a verification standard calibrated against samples of their own work.

The curriculum

13 modules in teaching order

Every module ends in a lab, and every lab produces an artifact your team keeps. Tracks take different paths through this list — nobody sits through all of it.

Curriculum
01

Portfolio triage and the scaling gap

Why enterprise-wide copilots scale easily and change nothing, while the function-specific work that moves the P&L stays stuck in pilot. One triage pass over what you already have running.

Horizontal vs vertical use casesAdoption depth vs adoption breadthScoring rubricBlocker diagnosis
In the lab2h

Sort every live AI initiative horizontal or vertical, score it, and name its blocker with evidence. Output is a ranked shortlist of three candidates that carries through every later module.

02

Workflow re-architecture

The spine of the programme. Factories that swapped steam engines for electric motors and kept the line-shaft layout got nothing; the gains came when the floor was redrawn. Same story here.

As-is mapping: swimlanes, handoffs, queue vs touch timeTask automation as a failure modeAgent-first redesignDecision rights
In the lab4h

Map one real end-to-end process using cycle times pulled from at least eight weeks of history, then redraw it agent-first and mark every point where a human still decides.

03

The capability boundary

Model capability does not track apparent difficulty. Tasks that look equally hard fall on opposite sides of the boundary, so it has to be discovered by testing rather than predicted by intuition.

The jagged frontier as a working modelDelegate / co-work / keepUncritical acceptanceVerification standards
In the lab3h

Blind-score a calibration harness built from your own work samples, spanning both sides of the boundary, then write the team’s delegation and verification standard.

04

Context engineering against your systems of record

Enterprise agents mostly fail on context and integration, not model quality. Conflicting metric definitions and undocumented conventions break more agents than weak reasoning does.

Business-definition ambiguityMCP connections to Jira, GitHub, Linear, NotionRetrieval and groundingKnowledge curation
In the lab2h

Populate a knowledge base with the definitions and conventions that currently live only in people’s heads, then ground the workflow in it and measure what changes.

05

Patterns before agents, and when not to use AI

Two refusals taught together: reaching for an autonomous agent when a deterministic three-step workflow ships this week, and reaching for AI when the correct answer is that the task should not use it.

Chaining, routing, parallelisation, orchestrator-workers, evaluator-optimiserSimplest pattern that worksSilent failure modesUnassignable accountability
In the lab3h

Build the deterministic version of your re-architected process, run it against ten real historical inputs, and record exactly where it holds and where it breaks.

06

Agents, orchestration and handoff contracts

Deliberately compressed. Most of multi-agent design is theatre; what is genuinely hard is specifying what each agent receives, returns, and is forbidden to assume.

Task decompositionOrchestrator and worker rolesHandoff contractsHuman checkpoints
In the lab2h

Write the handoff contract for every edge of a multi-agent version of your workflow, then build only the one handoff the team argues is genuinely load-bearing.

07

Evals and measurement with a counterfactual

An eval set that tells you whether the system is good, and a measurement design a finance partner will accept. A two-hour before-and-after comparison is not evidence.

Eval sets from real historical casesRubrics and model-as-judge gradingRegression testing across model upgradesHoldouts and staged rollout
In the lab3h

Assemble at least twenty real historical cases, write the grading rubric, run the eval, set the pass threshold, and write the one-page measurement design.

08

Cost and unit economics

The question a CFO asks first and most curricula never answer: what does one run cost, where is the ceiling, and how does that compare with the copilot licences already on the invoice.

Per-run and per-workflow unit costRate limits and quota planningCost regression on model upgradesBuild vs licence comparison
In the lab2h

Instrument the workflow for cost against real historical volume, compute cost per run and projected monthly cost at full rollout, then set the ceiling and the alert.

09

Indirect prompt injection and the agent threat model

An agent that reads Jira and writes to GitHub is a textbook injection target, and detection is not a reliable control. Capability scoping is the primary defence.

Injection through retrieved contentMCP servers as supply-chain dependenciesConfused-deputy problemsEgress allowlisting and sandboxing
In the lab3h

Red-team your own workflow: attempt injection through a connected system of record, a permission escalation, an unauthorised write, and exfiltration through a tool call.

10

Data protection and regulatory readiness

The module that decides whether legal lets any of this happen. Delivered in EU and non-EU variants, because AI Act obligations are real in one and a waste of half a day in the other.

EU AI Act Article 4 literacy obligationAnnex III classification and deployer dutiesGDPR Article 35 DPIAPrompt-input data classification and retention
In the lab3h

Classify your three shortlisted candidates, draft the DPIA for the one going to build, and write the rules for what may and may not enter a prompt.

11

Production operations and agent reliability

Where vertical use cases actually die. Including the question nobody asks until the first bad night: what is the rollback path once the agent has already closed the ticket and pushed the branch.

Signature-verified event triggersQueued execution, retries, idempotency, dead lettersCompensation and rollbackTrace review and on-call ownership
In the lab3h

Wire a real trigger from Jira or GitHub, run in shadow mode against live events, review the traces, then write the permission matrix and the runbook.

12

Freed capacity and the redeployment decision

Time saved with no plan for it evaporates. This is a decision pack, not a signing session — and where works councils apply, that conversation has an owner who is not in the room.

Measuring released capacity rather than assuming itThroughput vs quality vs coverage vs costMetric and decision-right consequencesConsultation obligations
In the lab1h

Produce the decision pack: measured capacity released, redeployment options with their metric implications, and the named forum where the choice actually gets made.

13

Handover and the recommendation pack

Close the loop without pretending a demo room can allocate budget. Teams show what runs, including the failures, and the sponsor books a real decision date.

Live demonstration against real inputsEval results and unit cost, failures includedSequencing the next two candidatesPlatform and residency decisions
In the lab1h

Demonstrate the running workflow, present results honestly, and assemble the recommendation pack the sponsor takes into the budget cycle.

13 modules · 32 contact hoursRequest the syllabus

Formats

Four ways to run it

Executive steering briefing

Half day + two checkpoints

Leadership teams who have funded AI, seen pilots, and cannot explain why none of it has reached the P&L.

  • Anonymised usage baseline vs what the room estimated
  • Portfolio triage against one rubric
  • Unit economics and regulatory exposure
  • Decision forum booked in the real budget cycle

Five-day workflow intensive

26 contact hours · max 10 people

One function with one process worth redesigning and a named owner who can actually change it.

  • Modules 01–08 and 11–13 built hands-on
  • Governance modules delivered as a briefed block
  • One workflow running in shadow mode by day five
  • Two facilitators throughout

Eight-week practitioner cohort

20 live hours · 2.5h per week

Rolling capability across several teams, when transfer into daily work matters more than speed.

  • Delegation, context, patterns and evals at double depth
  • Protected lab time between sessions
  • Each participant ships a workflow they use weekly
  • Governance covered as briefings

Four-week embedded sprint

Four weeks alongside your team

A candidate workflow stalled on integration, permissions, injection risk or legal sign-off.

  • Self-hosted instance stood up in your environment
  • MCP connections to your systems of record
  • Threat model with capability scoping enforced
  • Handover to your platform and security teams

What we will not tell you

The parts most proposals leave out

01

We will not quote you an ROI, EBIT or productivity figure, and we will decline to put one in a proposal. The research this curriculum is built on is largely self-reported and correlational, and we say so in the room.

02

We publish what we have not built. Open Backlog ships Jira, GitHub, Linear and Notion connectors — there is no Confluence or Bitbucket connector today. If those are your systems of record, we scope an export path or we decline the engagement.

03

Hardware isolation contains untrusted code execution inside your infrastructure. It does not stop prompt contents reaching a third-party inference API. Data residency is a separate question and we answer it in writing.

04

Open Backlog is AGPL-3.0 and enterprise legal frequently refuses network copyleft. A commercial licence exists. We would rather have that argument in week one than in procurement.

Research behind the curriculum

Every module traces to published evidence

We build the curriculum from field experiments and survey research rather than vendor decks, and we tell participants which findings are causal and which are self-reported and correlational.

Start with one workflow

Tell us which process costs you the most and which tools it runs through. We will come back with a track, a format and an honest read on whether this is worth doing at all.