# Alexander C.S. Hendorf — Full Profile > Independent AI & Open Source Strategist at opotoc GmbH — Solo Consultant Website: https://hendorf.com Languages: German, English Location: Germany Contact: https://hendorf.com/en/#contact --- ## About Alexander C.S. Hendorf is an independent AI and open-source strategist with more than 20 years of experience. He works as a solo consultant — personally, without a team or account managers — providing full accountability from one strategist. He specialises in regulated industries, particularly financial services, helping organisations navigate AI strategy, vendor evaluation, supplier oversight, and regulatory compliance (DORA, MaRisk, ESG). A Fellow of both the Python Software Foundation and the EuroPython Society, he leads the Open Source Working Group at the KI-Bundesverband and serves on the board of the Python Software Verband. His hands-on experience spans 50+ technologies including NLP, vector databases, workflow orchestration, and the broader Python ecosystem. Open source did not become relevant to me because of ideology, but because of responsibility. More than 20 years ago, I saw firsthand what technological dependencies cost companies. As COO of a transatlantic music company that I helped build, I learned that resilient structures emerge where companies retain control over their own ability to execute. Since then, I have worked at the intersection of open source, AI, and business transformation — initially as a developer, then as an architect, and today as a strategist to companies in regulated industries. When I recommend a technology, I know it from practice: the Python ecosystem, vector databases, workflow orchestration. I bring together strategic perspective and technical substance — from privacy-compliant architecture and IT security to resilient AI systems. When AI generates code or acts autonomously, quality assurance becomes the critical function. Hands-on depth makes reviews substantive, not nominal. Practice is based in Heidelberg, Germany. Engagements across DACH and Europe — on-site with clients or remote, depending on phase and confidentiality. ### Credentials & Affiliations - German AI Association | Head of Open Source Working Group - Python Software Foundation Fellow - Pioneers Hub Initiator - PySV Board Member - 50+ technologies deployed across enterprise and open-source ecosystems - 100+ conference talks in 15+ countries (ESA, UN AI for Good Geneva, WHU, PyData, EuroPython) - Featured in Handelsblatt, ZEIT, Tagesspiegel - Company: opotoc GmbH --- ## Services ### 01. AI Strategy with Execution Foundation For decision-makers who need robust clarity before committing to an investment. - A prioritised AI roadmap — what creates value, what waits, what stays out. - Independent assessment of vendors, platforms, and agents — with no commercial bias. ### 02. Execution Governance for AI Transformation For leaders who need to de-risk execution. I steer, your team delivers. - I secure architecture decisions, review vendor proposals, and keep execution on track. - My goal is capability, not dependency — your team moves forward confidently after the mandate. ### 03. Interim Leadership Mandate For organisations that need operational AI or data leadership on an interim basis. - I take leadership responsibility when AI or data structures need to be built, stabilised, or bridged. - Titles are secondary. What matters is that the function is filled effectively. Page: https://hendorf.com/en/interim-ai-leadership/ --- ## Current Research: The Governance Gap Most companies do not have an AI-governance problem. They have an unresolved architecture of accountability. Unresolved means: responsibility for AI shifts in four forms at once. Legal liability, internal accountability, operational responsibility, sign-off. ### Survey Findings - 71%: say accountability for AI-caused defects is unregulated and lands on the developer - 40%: name unclear governance as the No. 1 bottleneck, ahead of technology - 6%: feel relieved at day's end; 67% experience denser, more demanding days Source: Python Software Association Germany · a survey of 383 software professionals · June 2026 ### The Boardroom-to-Code Framework - Board: strategy, risk, ROI - Legal and compliance: liability, EU AI Act - Architecture: ownership, review, evaluation - Code: engineering's daily reality ### Three-Stage Model 01. The diagnostic (Stage 1: the risks are prioritised): I read your organisation across every level, from the boardroom to the code, along the framework. It shows where responsibility is unclear and where risk concentrates. 02. The Boardroom-to-Code Session (Stage 2: the board can decide): A focused session with your leadership team. Priorities, responsibilities and trade-offs get decided. It ends in the Executive Briefing. 03. The operating model (Stage 3: ownership becomes dependable): The design of the operating model: ownership and sign-off, review and eval standard, connected to your legal and compliance function, EU AI Act included. Documented board-ready. ### Five Dimensions of Unassigned Responsibility - Accountability: Who signs off AI code, and does that hold up in an audit? - ROI reality: Does the speed still pay after review, rework and risk? - Value migration: Where does the value of your engineering move, and does the role model follow? - Deskilling: In three years, will your organisation still understand its own code? - Shadow AI: Who decides on tools that are already in use? Full offering: https://hendorf.com/en/boardroom-to-code-s26/ Decision paper (PDF, delivered by email after a short form): Read first, talk after: the decision paper. --- ## What Sets Alexander Apart - **Foundation over short-termism**: I do not chase short-term effects. I work toward the technological and organisational foundations on which resilient AI capability is built: with clear standards, viable architecture, and the discipline to leave out what does not matter. - **From boardroom to code**: I move credibly between leadership, business functions, and engineering because I do not just present strategy — I assess it technically and translate it all the way into architecture decisions. - **Sovereignty is an architectural question**: Models are obsolete in months, vendors consolidate in quarters. Sovereignty is not a question of model choice, but of architecture: data-flow boundaries, vendor decoupling, auditability, exit paths. A resilient architecture lasts for years — and decides whether your company stays in command or follows a platform. - **Independent by design**: No software, no licences, no commissions. My recommendations follow your situation, not my revenue. --- ## Case Studies ### Quantitative Asset Management **Challenge:** Legacy C# and SAS silos separated research, portfolio management, and engineering. Long release cycles, limited ESG capability, and a lack of traceability put speed, control, and compliance at risk. **Result:** Migration of the organisation to Python and open source. Deployment cycles dropped from three months to three weeks; 44 employees were upskilled across four cohorts. The result was a stack that supports traceability, auditability, and regulatory requirements in a financial-services environment. Read more: https://hendorf.com/en/case-studies/transformation-asset-management ### Public Infrastructure **Challenge:** More than 30 years of project data, 70% manual processing, and no reliable data foundation. Publicly funded information was trapped in silos; AI potential could not be put to operational use. **Result:** Development of a data strategy with a 120-page implementation roadmap. An AI-supported forecasting model for workforce planning achieved 90% accuracy; the planning cycle shifted from annual to continuous. Read more: https://hendorf.com/en/case-studies/data-strategy-infrastructure ### Automotive / Research and Development **Challenge:** More than 10,000 research documents spanning three decades: unstructured, confidential, and nearly impossible to search. Cloud and API usage were ruled out. **Result:** Built an NLP-based knowledge explorer that turned static document storage into search results in seconds. Topics, related contexts, and source documents became directly accessible — fully on-premises, with no external APIs and no LLMs. The internal team subsequently expanded the solution further. Read more: https://hendorf.com/en/case-studies/nlp-knowledge-extraction --- ## Glossary Curated AI glossary for decision-makers in regulated industries. Each entry: definition, vendor-narrative vs. technical/regulatory reality, and the right boardroom question. ### Agentic AI URL: https://hendorf.com/en/glossary/agentic-ai/ **Short definition:** AI systems that autonomously plan and execute multi-step tasks across multiple tool calls. **Definition:** AI systems that autonomously plan and execute multi-step tasks across multiple tool calls — typically built on top of a large language model that orchestrates external interfaces, data sources and tools. **Noise — Signal:** In vendor decks, agentic AI appears as "AI now takes over what employees previously had to do manually". In reality, the production systems of 2026 are tightly bounded workflows with defined tools, deterministic guardrails and clear escalation logic. Open-acting agents — the kind that pick the task, the tools and the order themselves — are practically not production-ready in regulated industries. What works looks less spectacular from the outside than the demo. **The right question:** Not: "Where can we deploy agents?" But: "Which of our processes have tightly defined steps with auditable outputs — and which of those steps justify the additional cost of agent architecture, tracing and governance over a classic workflow automation?" ### AI Act (EU AI Act) URL: https://hendorf.com/en/glossary/ai-act/ **Short definition:** EU Regulation 2024/1689 governing AI systems, with staggered application through 2027 and separate obligations for general-purpose AI models. **Definition:** Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence, in force since 1 August 2024, with staggered application through 2027. Classifies AI systems by risk (prohibited, high-risk, limited, minimal) and contains separate obligations for general-purpose AI models. **Noise — Signal:** The AI Act is often reduced in the market to "we have to keep an AI inventory". In reality it bites at three very different levels: prohibitions of certain practices (since February 2025), obligations for high-risk systems (conformity assessment, logging, human oversight, transparency duties) and separate requirements for GPAI model providers and users. Which level affects a company doesn't depend on the AI tool it uses, but on the application context. **The right question:** Not: "Are we AI Act compliant?" But: "In which role — provider, deployer, importer, distributor — do we appear for which system, into which risk class does the use case fall, and which obligations take effect from which date?" ### AI Red Teaming URL: https://hendorf.com/en/glossary/ai-red-teaming/ **Short definition:** Structured, adversarial testing of an AI system against security, bias, hallucination and misuse patterns. **Definition:** Structured, adversarial testing of an AI system by dedicated teams or tools that deliberately surface security, bias, hallucination and misuse patterns. Distinct from classical software pentesting in its focus on model-inherent risks — prompt injection, jailbreaks, harmful outputs, data leakage through generative responses. **Noise — Signal:** Red teaming gets sold as "we're testing the AI for security". The phrase obscures a decisive distinction: provider red teaming before model release tests different risks than deployer red teaming before an application goes live. Most security gaps in productive AI applications arise from the specific combination of model, tools, permissions and data — and are not covered by provider red teaming. Outsourcing compliance to the model provider shifts a responsibility that cannot, in fact, be handed off. **The right question:** Not: "Has the provider red-teamed the model?" But: "Which attack vectors arise from the specific combination of our application — tools, data access, permissions, workflows — and who tests that combination before it goes into production?" ### Context Window URL: https://hendorf.com/en/glossary/context-window/ **Short definition:** Maximum number of tokens a language model can process simultaneously per request. **Definition:** Maximum number of tokens (roughly: word fragments) a language model can process simultaneously per request — including input, retrieved documents and output. In 2026, values range from around 8,000 (small models) to over two million tokens (frontier models). **Noise — Signal:** Larger context windows are marketed as "the model can now read all of our documentation". Technically true. In practice the model does not use the long context evenly — content at the beginning and end is processed more reliably than content in the middle ("lost in the middle" effect). Inference cost also scales linearly or super-linearly with context length, which can make large requests uneconomic. **The right question:** Not: "Do we need the model with the largest context window?" But: "What is the typical input size of our use cases, how do we measure whether the model actually finds the relevant passages, and when is a RAG architecture more economical than a larger context window?" ### Differential Privacy URL: https://hendorf.com/en/glossary/differential-privacy/ **Short definition:** A mathematical privacy standard that bounds the influence of any single record on a computation by a quantified parameter (epsilon). **Definition:** A mathematical privacy standard that guarantees that adding or removing any single record changes the result of a computation only by a clearly quantified amount (epsilon, ε). Implemented by adding controlled noise to data, intermediate results or model updates. **Noise — Signal:** Differential privacy is cited as "the gold standard for data protection" — which is correct, but rarely understood in the way the term is meant. The protection is parameterised: a high epsilon offers little protection and high data utility; a low epsilon offers strong protection but heavily noised results. In practice, the epsilon values used in vendor products are often so large that the protection against real adversaries is marginal. DP also doesn't defend against every threat — against an attacker with background knowledge of the data structure, the guarantees are weaker than communicated. **The right question:** Not: "Are we using differential privacy?" But: "At what epsilon, against which threat model, with what measured impact on data utility — and which additional measures complement DP, because on its own it isn't enough?" ### Distillation URL: https://hendorf.com/en/glossary/distillation/ **Short definition:** A technique for training a smaller model to approximate the behaviour of a larger one — at significantly lower inference cost. **Definition:** A technique for training a smaller "student" model to approximate the behaviour of a larger "teacher" model. Goal: comparable quality at significantly lower inference cost and latency. Frequently used in combination with fine-tuning on domain-specific tasks. **Noise — Signal:** Distillation is often simplified to "we just take a smaller model and become cheaper". Success depends on three conditions: high-quality training data or an accessible teacher model, a tightly defined task, and systematic evaluation. Distillation without a clearly bounded use case produces a model that looks stable on benchmarks but breaks at the edge cases. **The right question:** Not: "Can we distil the model to save costs?" But: "Which sub-task is tightly enough defined that a smaller model can map it reliably, and do we have the evaluation infrastructure to detect quality drift before it reaches the end customer?" ### DORA (Digital Operational Resilience Act) URL: https://hendorf.com/en/glossary/dora/ **Short definition:** EU Regulation 2022/2554 on the digital operational resilience of financial entities, applicable since 17 January 2025. **Definition:** Regulation (EU) 2022/2554, applicable since 17 January 2025. Governs the digital operational resilience of financial entities — ICT risk management, incident reporting, resilience testing and third-party risk management, including for cloud and AI service providers. **Noise — Signal:** In the AI context, DORA is often reduced to "we have to list our cloud providers". The substantive lever sits in third-party risk management (Article 28 ff.): critical ICT third-party providers — and that includes foundation-model providers as soon as they are embedded in business-critical processes — must be contractually auditable, must have documented exit strategies and must be included in resilience tests. Standard contracts from large US model providers rarely meet this bar today. **The right question:** Not: "Are our AI providers DORA-compliant?" But: "Which AI components in our value chains qualify as critical ICT functions, what does that mean contractually for audit rights and exit, and where does that force us toward open-source or on-premises alternatives?" ### Edge AI URL: https://hendorf.com/en/glossary/edge-ai/ **Short definition:** Running AI models directly on end devices instead of in central data centres or cloud APIs. **Definition:** Running AI models directly on devices at the edge of the network — smartphones, IoT devices, industrial sensors, vehicles — instead of in central data centres or cloud APIs. Typically uses quantised or distilled models to address hardware constraints on memory, compute and energy. **Noise — Signal:** Edge AI is often sold as "a privacy solution because nothing goes to the cloud". The privacy advantages are real but limited: the model itself comes from a cloud training environment, updates require telemetry, and the threat model shifts — device theft and reverse engineering become more relevant than API eavesdropping. Edge hardware is also the limiting factor: what runs on a smartphone in 2026 is significantly smaller than a typical frontier model, with correspondingly limited quality in language understanding, reasoning and multimodality. **The right question:** Not: "Can we push this to the edge?" But: "Which concrete requirements — latency, offline capability, privacy, bandwidth cost, regulatory locality — justify the edge stack, and which quality trade-offs do we accept against a cloud model for exactly those requirements?" ### Embedding URL: https://hendorf.com/en/glossary/embedding/ **Short definition:** A numeric vector representation of a text, image or data object in a high-dimensional space, where semantically similar objects sit close together. **Definition:** A numeric vector representation of a text, image or other data object in a high-dimensional space, where semantically similar objects sit close together. The basis for semantic search, classification, clustering and RAG. **Noise — Signal:** Embeddings are sold as "AI now understands meaning instead of just words". What they actually capture is the distribution of terms in the training corpus of the embedding model — not meaning in any epistemological sense. An embedding model trained on English web data performs worse on German technical text; a model from 2023 doesn't know the terminology of a 2025 regulatory text. The choice of embedding model is an architectural decision with consequences for search quality and compliance. **The right question:** Not: "Which embedding model is state of the art?" But: "On which language, which domain and which time period is the model trained, and how do we measure whether the search quality on our actual content — not on benchmarks — meets the requirements?" ### Evaluation (Eval) URL: https://hendorf.com/en/glossary/evaluation/ **Short definition:** Systematic, reproducible measurement of an AI system's quality against defined criteria. **Definition:** Systematic, reproducible measurement of an AI system's quality against defined criteria — typically using test datasets, metrics (accuracy, F1, BLEU, domain-specific scores), human-in-the-loop evaluation or LLM-as-judge approaches. **Noise — Signal:** "We tested the model" and "we evaluated the model" are not the same thing. Tests check whether the system runs. Evaluation measures whether it does the right thing — continuously, with documented datasets, defined metrics and thresholds at which a model is rolled back. The majority of AI initiatives that fail in production have no real eval infrastructure, because it was prioritised as "later". It does not arrive later. **The right question:** Not: "Does our model work?" But: "Which eval datasets, metrics and acceptance thresholds did we define before go-live, who checks them continuously, and what is the trigger for rollback or model change?" ### Federated Learning URL: https://hendorf.com/en/glossary/federated-learning/ **Short definition:** A training approach in which a shared model is trained across multiple decentralised data sources without the training data leaving those sources. **Definition:** A training approach in which a shared model is trained across multiple decentralised data sources without the training data itself leaving those sources. Instead, only model updates — gradients or weight deltas — are exchanged and aggregated on a coordinating server. **Noise — Signal:** Federated learning is marketed as a "privacy miracle": data stays local and the model still gets better. The reality is more nuanced. First, model updates can under certain conditions be analysed in ways that allow training data to be partially reconstructed (membership inference, gradient leakage); real privacy additionally requires differential privacy or secure aggregation. Second, federated setups significantly increase complexity in orchestration, versioning and evaluation. Third, the approach works best when the distributed data sources are homogeneous — and the heterogeneity that actually justifies federated learning is at the same time its biggest quality lever. **The right question:** Not: "Can we use federated learning to protect our data?" But: "Which regulatory or contractual requirement actually forbids merging the data, which additional privacy mechanisms are needed for it, and is the complexity overhead in proportion to the actual privacy improvement over centralised training with DP?" ### Fine-Tuning URL: https://hendorf.com/en/glossary/fine-tuning/ **Short definition:** A technique for retraining the weights of a pre-trained foundation model on a domain-specific dataset. **Definition:** A technique for retraining the weights of a pre-trained foundation model on a domain-specific dataset to improve performance on a narrower class of tasks. Variants range from full fine-tuning through parameter-efficient methods (LoRA, QLoRA) to instruction tuning and reinforcement learning from human feedback. **Noise — Signal:** Fine-tuning is recommended as the default answer to "the model doesn't know our data". In most cases it is the worse choice. For up-to-date knowledge, RAG is faster, cheaper and more auditable; for behavioural steering, prompting plus few-shot is usually enough; fine-tuning justifies itself primarily for stable, tightly defined tasks with sufficiently high-quality training data and a clear evaluation strategy. Fine-tuning without these preconditions produces models that are more expensive than the original and slowly degrade in production. **The right question:** Not: "Should we fine-tune the model on our data?" But: "What is the specific task, what training data and eval infrastructure do we have, and which of the alternatives — RAG, prompting, a smaller model with a better pipeline — have we examined before choosing fine-tuning?" ### Foundation Model URL: https://hendorf.com/en/glossary/foundation-model/ **Short definition:** A large, generally pre-trained AI model that serves as the basis for a variety of downstream applications. **Definition:** A large, generally pre-trained AI model that is fine-tuned or adapted via prompts as the basis for a variety of downstream applications. The term covers language models, multimodal models and domain-specific pre-training; it is broader than "large language model". **Noise — Signal:** "Foundation model" and "LLM" are used synonymously, but they are not the same — every LLM is a foundation model, but not the other way around. More important in the boardroom is the legal dimension: in the EU AI Act, "general-purpose AI model" has a very specific meaning with thresholds above which transparency, documentation and risk-management obligations kick in — also for companies that deploy such a model without training it themselves. Anyone who buys, integrates or further develops foundation models is not automatically only a "user". **The right question:** Not: "Which foundation model are we deploying?" But: "In which role do we appear — user, deployer, provider, modifier — and which AI Act obligations follow if we fine-tune the model or enrich it with our own data?" ### Guardrails URL: https://hendorf.com/en/glossary/guardrails/ **Short definition:** Mechanisms before, during or after model inference that filter, restrict or escalate undesired inputs or outputs. **Definition:** Mechanisms before, during or after model inference that filter, restrict or escalate undesired inputs or outputs — from simple regex filters through classifiers to specialised guardrail models and policy engines. **Noise — Signal:** Guardrails are often sold as "the safety layer that makes the model safe". They are not. They are a layer of additional heuristics between input and model, or between model and user. They reduce risk, they do not eliminate it, and they have their own failure modes: false positives that block legitimate requests, false negatives that let problematic content through. In regulated industries, guardrails do not replace risk management — they are one building block within it. **The right question:** Not: "Which guardrails do we need?" But: "Which risks do we address at which layer — input, model, output, workflow — how do we measure the hit rate in live operation, and which risks remain structurally outside what guardrails can deliver?" ### Hallucination URL: https://hendorf.com/en/glossary/hallucination/ **Short definition:** Output of a language model that is plausibly worded but factually wrong or not supported by the sources. **Definition:** Output of a language model that is plausibly worded but factually wrong or not supported by the provided sources. Hallucinations are not bugs in any narrow sense — they are a structural property of probabilistic language models, which predict word sequences, not truth. **Noise — Signal:** "Hallucination" suggests a pathological special case that can be fixed. In fact, hallucination is the default mode of a language model; what appears as a correct answer is a hallucination that, by chance or via suitable context, happens to coincide with reality. Consequence: "we reduce hallucinations" is the right phrasing; "we prevent hallucinations" is not. Techniques like RAG, constrained decoding or chain-of-thought lower the rate, they do not eliminate it. **The right question:** Not: "How do we prevent hallucinations?" But: "At which points in our application path is a wrong statement tolerable, at which not, and which verification layer — source attribution, human sign-off, external fact-checking — kicks in before the result reaches the end customer?" ### Harness URL: https://hendorf.com/en/glossary/harness/ **Short definition:** A structured software layer around an AI model that orchestrates tool calls, eval routines, guardrails and output processing. **Definition:** A structured software layer that connects an AI model to the outside world — through defined inputs, tool calls, eval routines, safety checks and output processing. In 2026 the term is used in two main contexts: as eval harness (systematic testing infrastructure such as lm-evaluation-harness or OpenAI Evals) and as agent harness (runtime scaffolding that turns a model into an agentic system — tool calls, memory, escalation paths). **Noise — Signal:** When vendors say "we deliver foundation model X", that's only half the story — quality and safety of a productive system come to a substantial degree from the surrounding harness. Two applications with identical models can deliver very different results depending on how the harness exposes tools, curates context, catches hallucinations and builds in eval loops. The common shortcut in the market: the model is bought, the harness is "somehow built ourselves". That is exactly where lock-in, security gaps and migration risk emerge. **The right question:** Not: "Which model are we deploying?" But: "Which components of our harness — tool routing, context curation, eval loop, guardrails, escalation — do we build ourselves, which do we buy, and where does lock-in arise if we want to swap the model?" ### Inference Cost / TCO URL: https://hendorf.com/en/glossary/inference-cost-tco/ **Short definition:** Ongoing cost of model use per request; TCO extends this to development, eval, hosting, monitoring and compliance over the lifecycle. **Definition:** Inference cost denotes the ongoing cost of model use per request — typically per token, per image or per second of compute. Total cost of ownership (TCO) additionally encompasses development, data preparation, eval, hosting, monitoring, retraining and compliance overhead across the lifecycle. **Noise — Signal:** AI business cases are routinely calculated on the basis of list prices ("$5 per million tokens"). In production, actual costs typically run three to ten times higher: long prompts, multiple model calls per user action (routing, reasoning, verification), retries, eval calls, monitoring pipelines. On top of that come infrastructure scaling costs, which are pushed onto the provider with foundation-model APIs but remain visible in on-premises setups. **The right question:** Not: "What does the model cost us?" But: "What are the full costs per productive user action across the entire application path — and how does the ratio change when we scale by a factor of 10 or 100?" ### Loop Engineering URL: https://hendorf.com/en/glossary/loop-engineering/ **Short definition:** Instead of prompting an AI agent step by step, you design the system that prompts it — one that finds, dispatches, checks and remembers work on its own. **Definition:** A practice that became popular in early 2026: rather than prompting a coding agent by hand, you build the system that does it for you. A loop is a recursive goal — you define a purpose, and the loop iterates until a verifiable condition holds. Five building blocks carry it. Automations set the cadence, kicking off runs on a schedule. Git worktrees isolate parallel agents so their edits cannot collide. Skills capture project knowledge that would otherwise be re-explained every session. Plugins and connectors via MCP wire the loop into the real tools — issue tracker, database, Slack, pull request. Sub-agents separate writing from checking: the agent that wrote the code is not the one that signs it off. A sixth element decides whether any of it works: a memory outside the conversation — a markdown file, a Linear board — because the model forgets everything between runs. Claude Code and Codex now ship all of these out of the box; commands such as `/loop` (re-runs on a cadence) and `/goal` (runs until an externally verified condition holds) move control one floor higher. The phrase traces to Peter Steinberger and to Boris Cherny, head of Claude Code at Anthropic ("my job is to write loops"). **Noise — Signal:** Loop engineering is being sold as the next phase after prompt engineering — the proof that a team has moved from holding the tool to designing the factory. The capability is real and probably a preview of how this work evolves. Three things, however, get sharper, not easier, as the loop improves. First, the token bill. A loop runs unattended — and bills unattended. Sub-agents multiply the cost because each one does its own model and tool work. The vendor incentive is not neutral here: more loops mean more tokens, and tokens are what foundation-model vendors monetise directly. Second, verification. A loop running unattended is a loop making mistakes unattended. "Done" is a claim the loop makes about itself — not a proof. Splitting the writing sub-agent from the checking one is the reason you can step away at all, not a hygiene option. Third, the human side. The same loop accelerates someone who understands the work deeply — and lets someone else avoid understanding it altogether. Identical setup, opposite outcomes. Loop design is therefore harder than prompt engineering, not easier: the leverage point moved, the work did not disappear. **The right question:** Not: "Can our teams build loops that run the agents for us?" But: Who verifies what the loop ships while it runs unattended? Is the stopping condition externally checkable? What does an unattended mistake cost — and how do we keep the token bill governable as sub-agents multiply? And is the loop being used to move faster on work people understand — or to avoid understanding it? ### Mixture of Experts (MoE) URL: https://hendorf.com/en/glossary/mixture-of-experts/ **Short definition:** An architectural pattern in which a model consists of several specialised subnetworks, only a small selection of which is activated per token. **Definition:** An architectural pattern in which a language model consists of several specialised "expert" subnetworks, only a small selection of which is activated per token (sparse activation). Allows significantly higher total parameter counts at comparable inference compute cost. 2026 examples: the Mixtral family, the DeepSeek-V3 line and several frontier models whose architecture isn't public but is presumed to be MoE. **Noise — Signal:** MoE is sold as "the architecture with which we scale efficiently". In fact it just shifts the scaling axis: less inference compute per token, but higher memory requirements (all experts must be loaded), greater complexity in routing, and harder hardware utilisation — especially in on-premises setups with limited GPU memory. From the user-side perspective the relevant point is: at comparable quality, MoE models can be cheaper per token as long as the hosting carries the memory overhead; on your own hardware that assumption is not a given. **The right question:** Not: "Should we deploy MoE models?" But: "What implications does the MoE architecture have for our hosting (GPU memory, utilisation), our latency requirements and availability on on-premises stacks compared to dense models of similar quality?" ### Model Card URL: https://hendorf.com/en/glossary/model-card/ **Short definition:** A structured document describing the training data, intended use, limitations, bias patterns and licence of an AI model. **Definition:** A structured document describing an AI model: training data and date, intended use, limitations, evaluation results, known bias patterns, licence and legal terms. Introduced in 2018 (Mitchell et al.), now part of regulatory requirements — including under the AI Act for high-risk systems and GPAI models. **Noise — Signal:** Model cards are marketed as "transparency", but in practice rarely deliver any. Providers often publish only the favourable benchmarks and avoid statements about training-data provenance, cutoff date, problematic failure modes or bias tests. A model card that does not disclose what data was used for training and where the model fails is a marketing document. **The right question:** Not: "Does the model have a model card?" But: "Which of the regulatorily relevant items for us — training-data provenance, cutoff, documented failure modes, bias tests, licence terms — are contained in the model card, and which gaps must we close before deployment?" ### Model Governance URL: https://hendorf.com/en/glossary/model-governance/ **Short definition:** Processes, roles and documentation that steer the lifecycle of an AI model in a company. **Definition:** Processes, roles and documentation that steer the lifecycle of an AI model in a company — from selection through validation, sign-off, monitoring and re-approval to decommissioning. Encompasses responsibilities, eval requirements, risk classification, versioning, audit trails and escalation paths. **Noise — Signal:** Model governance is often summarised as an "AI governance framework" on a single slide — typically a diagram with arrows between roles such as "AI Lead", "Risk Owner", "Compliance Officer". Substantive governance isn't the diagram, it's the answer to concrete questions: who is allowed to deploy which model for which use case? At which risk level does which escalation kick in? Which eval thresholds are binding? Who monitors model drift in live operation, and what triggers a rollback? Without documented answers to these questions, model governance is a PowerPoint artefact that does not carry weight in audits. **The right question:** Not: "Do we need a model governance framework?" But: "Which of the concrete lifecycle decisions — model selection, sign-off, drift detection, re-approval, decommissioning — are today assigned by name to a role, with documented thresholds and a traceable decision history?" ### Multimodal URL: https://hendorf.com/en/glossary/multimodal/ **Short definition:** An AI model's capability to process multiple input and output modalities — typically text, image, audio and video. **Definition:** An AI model's capability to process multiple input and output modalities — typically combinations of text, image, audio and video. In 2026, multimodality is standard in frontier models and remains an architectural decision in specialised models. **Noise — Signal:** Multimodality is often presented as a universal capability. In practice the modalities are unevenly covered: text and image understanding are stable, audio generation is delicate in regulated applications (voice cloning, authenticity), video generation remains quality-sensitive. Cost and latency also scale with modality — an image in the prompt is often equivalent to several thousand tokens, video to several hundred thousand. **The right question:** Not: "Do we need a multimodal model?" But: "Which modality delivers demonstrable value over a text pipeline for which concrete use case, and does that value justify the cost and compliance implications?" ### On-Premises AI URL: https://hendorf.com/en/glossary/on-premises-ai/ **Short definition:** Operating AI models and infrastructure in your own or a dedicated environment instead of through the model providers' API services. **Definition:** Operating AI models and infrastructure in your own or a dedicated, model-provider-independent environment — typically in your own data centre or in a sovereignly controlled hyperscaler setup. Generally requires open-weights models and your own MLOps infrastructure. **Noise — Signal:** On-premises is often dismissed wholesale as "expensive and slow", or, in the opposite direction, glorified as "the only safe solution". Both miss the point. The question isn't cloud vs. on-premises but: which workloads, given their regulatory status, data sensitivity or volume profile, justify the effort of running your own stack? For high-frequency, latency-critical, privacy-sensitive inference, on-premises is often the economically and regulatorily superior option — for sporadic knowledge-work applications, rarely. **The right question:** Not: "Should we go on-premises?" But: "Which of our AI workloads cross the thresholds at which on-premises pays — data sensitivity, regulatory status, volume, latency — and which open-weights models qualify in quality and licence for those workloads?" ### Open Weights vs. Open Source URL: https://hendorf.com/en/glossary/open-weights-vs-open-source/ **Short definition:** Open weights denotes publication of the model parameters; open source additionally requires training code, data specification and an open licence. **Definition:** Open weights means a provider makes the trained model parameters available for download. Open source additionally means that training code, data specification and a licence are available that explicitly permit redistribution, modification and commercial use — in the spirit of the Open Source Initiative's definitions and the Open Source AI Definition (OSAID). **Noise — Signal:** When a vendor advertises "open-source AI", they almost always mean open weights with a restricted licence. This isn't a semantic detail. It decides whether a model is auditable, can be operated on-premises, may be fine-tuned, and whether the supply chain can be documented for DORA, MaRisk and the AI Act. A licence that excludes commercial use above a certain size or restricts use cases is not open source — even if the model is marketed that way. **The right question:** Not: "Is the model open source?" But: "Which components — weights, code, training data, licence terms — are available under which conditions, and is that enough for our regulatory evidence obligations and our exit strategy?" ### Prompt Injection URL: https://hendorf.com/en/glossary/prompt-injection/ **Short definition:** An attack technique in which inputs are crafted so that the model ignores or overrides its original system instructions. **Definition:** An attack technique in which an attacker crafts inputs so that the model ignores, overrides or turns its original system instructions against the operator. Direct (in the user input) or indirect (in content the model processes — documents, web content, tool outputs). **Noise — Signal:** Prompt injection is often dismissed as something that "can be solved with a few filters". Today it is the OWASP top-1 risk factor for LLM applications, and there is no complete technical mitigation. Indirect prompt injection — instructions an attacker embeds in a document or email that the model later processes — is particularly relevant to agentic AI architectures and enterprise search. An application that processes external, untrusted content and is at the same time allowed to perform privileged actions is structurally vulnerable. **The right question:** Not: "How do we prevent prompt injection?" But: "At which points does our system process untrusted content, which actions is the model allowed to trigger at those points, and which permission design reduces the blast radius in the event of a successful injection?" ### RAG (Retrieval-Augmented Generation) URL: https://hendorf.com/en/glossary/rag/ **Short definition:** An architectural pattern in which a language model retrieves relevant documents from a knowledge base and integrates them as context in the prompt. **Definition:** An architectural pattern in which a language model, when answering a query, first retrieves relevant documents from a knowledge base — typically through semantic search in a vector database — and integrates them as context in the prompt. **Noise — Signal:** RAG is often sold as the solution to three separate problems: freshness, hallucination and compliance. It reliably solves only the first. Hallucinations are reduced, not eliminated — the model can misinterpret retrieved sources, mix contradictory ones or simply ignore the context. And compliance is a property of the overall system of source rights, audit trails, reasoning chains and access control, not of the architecture. A RAG pipeline alone fulfils no regulatory requirement. **The right question:** Not: "Should we introduce RAG to avoid hallucinations?" But: "How do we measure the quality of the retrieved sources, how do we document the reasoning chain of an answer, and who is accountable when the model gets the right context and still uses it wrong?" ### Reasoning Model URL: https://hendorf.com/en/glossary/reasoning-model/ **Short definition:** A class of language models that produce a longer internal chain of thought before answering and outperform classical LLMs on multi-step tasks. **Definition:** A class of language models, established since 2024/2025, that produce a longer internal chain of thought before the actual answer — visible or hidden — and deliver significantly better results on mathematical, planning and multi-step tasks than classical LLMs. 2026 representatives include the OpenAI o-series, Claude with Extended Thinking and the DeepSeek R1 line. **Noise — Signal:** Reasoning models are advertised as "AI that finally understands what it's doing". What they actually do is produce structured token sequences before the answer. That noticeably lifts quality on certain tasks but drives inference cost and latency by a factor of five to twenty. For most productive applications — classification, extraction, summarisation, routing — reasoning models are the more expensive and slower choice with no meaningful quality gain. They are a tool, not a default. **The right question:** Not: "Should we switch to reasoning models?" But: "What share of our application mix actually requires multi-step reasoning, what is better served by a fast, cheap model, and how do we steer the routing between the two?" ### Sovereign AI URL: https://hendorf.com/en/glossary/sovereign-ai/ **Short definition:** AI infrastructure under national or European control — across data, operations, models and training data. **Definition:** A politically loaded umbrella term for AI infrastructure under national or European control — across the layers of data, operations, models and training data. There is no single technical definition. **Noise — Signal:** "Sovereign" is used by vendors in at least three very different meanings: first as data residency (the data stays in the EU), second as hosting sovereignty (the infrastructure is operated by a European provider), third as full control over model, weights and training data. Only the third is real sovereignty. The first two are useful but not sovereign if model updates, licence terms or the training pipeline remain under the control of a non-European actor. This distinction is rarely flagged in vendor decks. **The right question:** Not: "Is this solution sovereign?" But: "On which of the three layers — data, operations, model — do we actually have control, which remains a dependency, and which of those dependencies is manageable in a crisis?" ### Synthetic Data URL: https://hendorf.com/en/glossary/synthetic-data/ **Short definition:** Artificially generated data that reproduces statistical or structural properties of real data. **Definition:** Artificially generated data that reproduces statistical or structural properties of real data — produced by language models, generative image models, simulations or rule-based methods. Used for training, augmentation, testing or privacy-compliant analytics. **Noise — Signal:** Synthetic data is sold as a way out of data scarcity and privacy problems. Both only partially true. Synthetically generated data inherits the bias of the generators and rarely captures edge cases that occur in real operation — it is a feedback amplifier for known patterns, not a generator for unknown ones. Under data-protection law it is treated as non-personal only if re-identification is demonstrably ruled out; many implementations don't deliver that proof. **The right question:** Not: "Can we train the model on synthetic data?" But: "Which gap in our real data set should the synthetic augmentation close, how do we validate that it does, and which privacy and audit evidence is needed for synthetic data to hold up regulatorily?" ### Tokenmaxxing URL: https://hendorf.com/en/glossary/tokenmaxxing/ **Short definition:** Status game in which engineers and teams compete on raw token consumption — often treated as a proxy for productivity rather than a measure of outcome. **Definition:** Cultural and operational pattern, popularised in early 2026, of maximising token throughput per engineer via long-running coding agents, parallel sub-agents and stacked subscriptions. Internal leaderboards at some AI labs rank workers by tokens consumed, "token budgets" are offered as a perk, and weekly footprints in the hundreds of millions to billions of tokens per person are no longer unusual. Some managers factor token use into performance reviews. **Noise — Signal:** Token consumption is being marketed — internally and externally — as evidence of AI adoption and engineering productivity. In reality it measures input, not output quality. Continuously running agents produce throwaway code, redundant retries, speculative refactors and expensive eval loops; a leaderboard rewards motion, not merit. The incentive is not neutral: foundation-model vendors monetise tokens directly, and Anthropic and OpenAI revenue growth in 2026 is driven largely by agentic coding token volume. "More tokens" is the metric the seller wants the buyer to optimise. **The right question:** Not: "How many tokens is our team consuming?" But: "What is our cost per shipped feature, what is the defect rate of agent-generated code in production, and how do we separate genuine productivity gains from leaderboard-driven theatre — before token budgets become a line item nobody can justify?" ### Vector Database URL: https://hendorf.com/en/glossary/vector-database/ **Short definition:** A database optimised for storage and search of high-dimensional vectors (embeddings). **Definition:** A database optimised for storage and search of high-dimensional vectors (embeddings). Core function: fast approximate nearest-neighbour search to retrieve semantically similar content. 2026 representatives: specialised engines (Qdrant, Weaviate, Milvus) and vector extensions of relational or document-oriented databases (e.g. pgvector for PostgreSQL). **Noise — Signal:** Vector databases are presented in RAG marketing as "the missing component". In fact the database is the technically least critical element in a RAG pipeline — the levers sit in chunking strategy, embedding model, hybrid search (vector plus keyword), re-ranking and source-rights management. A poorly designed pipeline on a state-of-the-art vector database delivers worse results than a cleanly built pipeline on PostgreSQL with pgvector. **The right question:** Not: "Which vector database do we need?" But: "What requirements do we have on volume, latency, filtering, hybrid search and multi-tenancy, and in which order do we optimise chunking, embedding and re-ranking before the choice of database becomes the bottleneck at all?" --- ## Common Challenges Addressed - **Strategy gap**: AI is on the agenda, but pilot projects are not turning into a scalable operating model. What is missing are priorities, architecture, and a credible target picture. - **Technology sprawl**: Business units and analytics teams are experimenting in parallel with different tools, platforms, and vendors. That increases complexity, cost, and governance risk. - **Lost in the hype**: The market produces new promises every day. The real question is not what is new, but what is economically, technically, and regulatorily viable in your context. - **AI agents without guardrails**: Guardrails, boundaries, and operational governance are often missing. Without robust safeguards, AI outputs create risk rather than value. - **Execution bottleneck**: The direction is clear, but business, technology, and leadership are not operating from the same logic. Decisions stall, and execution slows down. - **Digitisation instead of transformation**: AI does not create impact by digitising existing routines, but by redesigning how decisions are made and how work gets done. --- ## Speaker & Keynotes 100+ talks on international stages. From the European Space Agency to WHU to the United Nations in Geneva. ### Current Topics ### Agentic AI: From hype to value How companies can use AI agents strategically — what works, what fails, and why buying into the buzz does not build a business. ### AI & open source in regulated industries How banks, insurers, and other regulated organisations can adopt AI and open-source technology without getting trapped by legacy systems, compliance anxiety, or organisational inertia. ### AI strategy beyond tools Why standards, structure, governance, and the right questions matter more than the latest platform — and how to align everyone from engineering teams to the C-suite. ### Selected Talks - 2026: Fail and Grow: Why Failure Stories Matter in AI Development and Innovation — AI for Good, United Nations, Geneva [highlighted] - 2026: CrAIzy Times - KI zwischen Hype und Realität - Keynote — Unzer AI Week - 2026: Fireside Chat with Sebastian Raschka: Stop Waiting, Start Shipping: Real-World Strategy for Open-Source LLMs — PyCon DE & PyData, Darmstadt [highlighted] - 2025: Structured Automation with Agentic AI — European Space Agency (ESA), Darmstadt [highlighted] - 2025: The Agent Hype Trap: Why Buying Into the Buzz Won't Build a Business — hei_Innovation, Heidelberg University [highlighted] - 2025: The Future Belongs to the Curious: Trends, Diversity & the Power of AI — Unzer, Munsbach (Luxembourg) [highlighted] - 2025: Beyond Agents: What AI Strategy Really Needs in 2025 — PyCon DE & PyData, Darmstadt - 2025: The Future Belongs to the Curious - Keynote — Lidl Data Science Conference, Neckarsulm - 2025: AI in Reality Fireside Chat: Enterprise AI & Open-Source Innovation Fireside chat with Ines Montani (spaCy), the CDO of Merck, and Dr. Alexander Beck (Quoniam) — PyCon DE & PyData, Darmstadt [highlighted] - 2025: NVIDIA GTC 2025 Insights — Hessian AI Academy, TU Darmstadt - 2025: Open-Source Business Fireside chat with Yann Lechelle (CEO, probabl.), Sylvain Corlay (CEO, QuantStack) — PyData, Cité des Sciences, Paris - 2024: Why AI Projects Fail — KI Navigator, Nuremberg - 2023: Solving Data Problems in Management Accounting — EuroPython, Prague - 2022: Lessons Learned About Data & AI at Enterprises and SMEs — PyData, London - 2022: AI and Data Science — Context and the Big Picture — Free University of Berlin - 2020: The Economics of Prediction Machines — Smart Production Network Forum, ABB, Mannheim - 2020: AI for Lawyers — German University of Administrative Sciences, Speyer - 2019: AI for Managers — Guest Lecture — WHU — Otto Beisheim School of Management, Vallendar [highlighted] - 2016: MongoDB Data Analysis & Visualization — MongoDB World, New York City [highlighted] - 2016: Agile Datenanalyse - der schnelle Weg zum Mehrwert — ZEW - Zentrum für Europäische Wirtschaftsforschung [highlighted] - 2014–2019: EuroPython — Speaker & Program Chair — Berlin, Bilbao, Basel, Rimini, Dublin, and others --- ## Public Presence & References - Despite AI - Experts Consider these IT Jobs Future-proof (German) — Handelsblatt: https://www.handelsblatt.com/karriere/berufe-diese-it-jobs-halten-experten-trotz-ki-fuer-zukunftssicher/100238210.html - AI sovereignty begins not with the model, but with the infrastructure (German) — cloudmagazin: https://www.cloudmagazin.com/2026/06/01/ki-souveraenitaet-infrastruktur-open-source-hendorf/ - The Impact of AI Agents on IT Jobs (German) — Computerwoche: https://www.computerwoche.de/article/4172250/diese-it-jobs-sind-am-starksten-von-ki-betroffen.html - Good online learning plus practice often beats the traditional lecture today (German) — Die Zeit: https://www.zeit.de/zeit-spezial/2022/01/data-scientist-beruf-studium-quereinstieg/ - AI for Managers — Guest lecture at WHU: https://www.youtube.com/watch?v=ItjElD4QqrY - Which skills will matter for a top job in the future (German) — Tagesspiegel: https://www.tagesspiegel.de/wirtschaft/karriere/karriere-im-it-bereich-welche-kenntnisse-fur-einen-top-job-kunftig-zahlen-13615987.html --- ## Social & Contact - Website: https://hendorf.com - LinkedIn: https://linkedin.com/in/hendorf - GitHub: https://github.com/alanderex - X: https://x.com/hendorf - Mastodon: https://mastodon.social/@hendorf - Advisory contact: https://hendorf.com/en/#contact - Speaker booking: https://hendorf.com/en/speaker/#contact --- ## Site Structure - Homepage (DE): https://hendorf.com/de/ - Homepage (EN): https://hendorf.com/en/ - Boardroom-to-Code Session (DE): https://hendorf.com/de/boardroom-to-code-s26/ - Boardroom-to-Code Session (EN): https://hendorf.com/en/boardroom-to-code-s26/ - Speaker page (DE): https://hendorf.com/de/speaker - Speaker page (EN): https://hendorf.com/en/speaker/ - Case Study (DE): https://hendorf.com/de/case-studies/transformation-asset-management/ - Case Study (EN): https://hendorf.com/en/case-studies/transformation-asset-management/ - Case Study (DE): https://hendorf.com/de/case-studies/datenstrategie-infrastruktur/ - Case Study (EN): https://hendorf.com/en/case-studies/data-strategy-infrastructure/ - Case Study (DE): https://hendorf.com/de/case-studies/nlp-wissensextraktion/ - Case Study (EN): https://hendorf.com/en/case-studies/nlp-knowledge-extraction/ - Blog (DE): https://hendorf.com/de/blog/ - Blog (EN): https://hendorf.com/en/blog/ - Glossary (EN): https://hendorf.com/en/glossary/ - Glossary entry (EN): https://hendorf.com/en/glossary/agentic-ai/ - Glossary entry (EN): https://hendorf.com/en/glossary/ai-act/ - Glossary entry (EN): https://hendorf.com/en/glossary/ai-red-teaming/ - Glossary entry (EN): https://hendorf.com/en/glossary/context-window/ - Glossary entry (EN): https://hendorf.com/en/glossary/differential-privacy/ - Glossary entry (EN): https://hendorf.com/en/glossary/distillation/ - Glossary entry (EN): https://hendorf.com/en/glossary/dora/ - Glossary entry (EN): https://hendorf.com/en/glossary/edge-ai/ - Glossary entry (EN): https://hendorf.com/en/glossary/embedding/ - Glossary entry (EN): https://hendorf.com/en/glossary/evaluation/ - Glossary entry (EN): https://hendorf.com/en/glossary/federated-learning/ - Glossary entry (EN): https://hendorf.com/en/glossary/fine-tuning/ - Glossary entry (EN): https://hendorf.com/en/glossary/foundation-model/ - Glossary entry (EN): https://hendorf.com/en/glossary/guardrails/ - Glossary entry (EN): https://hendorf.com/en/glossary/hallucination/ - Glossary entry (EN): https://hendorf.com/en/glossary/harness/ - Glossary entry (EN): https://hendorf.com/en/glossary/inference-cost-tco/ - Glossary entry (EN): https://hendorf.com/en/glossary/loop-engineering/ - Glossary entry (EN): https://hendorf.com/en/glossary/mixture-of-experts/ - Glossary entry (EN): https://hendorf.com/en/glossary/model-card/ - Glossary entry (EN): https://hendorf.com/en/glossary/model-governance/ - Glossary entry (EN): https://hendorf.com/en/glossary/multimodal/ - Glossary entry (EN): https://hendorf.com/en/glossary/on-premises-ai/ - Glossary entry (EN): https://hendorf.com/en/glossary/open-weights-vs-open-source/ - Glossary entry (EN): https://hendorf.com/en/glossary/prompt-injection/ - Glossary entry (EN): https://hendorf.com/en/glossary/rag/ - Glossary entry (EN): https://hendorf.com/en/glossary/reasoning-model/ - Glossary entry (EN): https://hendorf.com/en/glossary/sovereign-ai/ - Glossary entry (EN): https://hendorf.com/en/glossary/synthetic-data/ - Glossary entry (EN): https://hendorf.com/en/glossary/tokenmaxxing/ - Glossary entry (EN): https://hendorf.com/en/glossary/vector-database/ - Imprint: https://hendorf.com/de/imprint/ - Privacy Policy: https://hendorf.com/de/data-policy/ --- ## Blog Posts (Full Text) ### Skill Shift, Not Replacement: What Developers Report About Working With AI URL: https://hendorf.com/en/blog/ai-developer-skill-shift-control-debt/ Published: 2026-07-27 Tags: ai-strategy, governance, enterprise-ai, ai-evaluation, future-of-work > Companies are automating production faster than they are building review and accountability. That gap is the new control debt. A great deal gets written about AI and developer jobs, mostly on the basis of studies from economics, consulting and HR. Through the Python Software Association Germany we took the opposite route and asked the people doing the work: 383 developers from the PyCon DE community reported how AI is changing their working day. Handelsblatt has reported on the findings. The strongest result is not about whether jobs disappear, which this survey does not measure, but about a structural gap. Intensive AI use creates new reviewing, accountability and assurance work for which the structures are often missing. I call that gap control debt. ## At a Glance - AI use is routine in this community: 96 percent reach for AI tools at least several times a day, and a good one in five automates more than half of their routine tasks. Whether that creates or removes jobs was not surveyed. - The work is moving rather than disappearing: 68 percent name reviewing and curating AI output as a new task, and only 9 percent see no significant new tasks at all. - The skill shift has a clear direction: domain knowledge (60 percent), architecture (56 percent) and requirements communication (48 percent) are gaining value, while AI is used most heavily for code, documentation and research. - The real gap is organisational: 71 percent say accountability for AI-induced defects is not explicitly defined, only 11 percent have a documented framework, and only 30 percent have a formal evaluation. That is exactly where control debt accumulates. ## We Asked Practitioners Instead of Talking About Them Most numbers on the future of developer work come from economic models, analyst reports or HR surveys that observe the people affected from the outside. This survey turns the perspective around and asks the people who put AI to productive use every day. The sample is self-selected and only partly representative of the German developer market in 2026. Respondents are predominantly experienced, actively invest in keeping their skills current and attend professional conferences: early adopters rather than the average. That is precisely what makes the survey valuable. It shows what is happening where AI is already used intensively. Their experience is an early signal of which changes could become relevant for other development teams as AI use spreads. ## The New Work: Reviewing, Curating, Steering The most robust finding sits on the other side of automation: AI takes tasks over and creates new ones at the same time. 68 percent of respondents name reviewing and curating AI output as a new task. 48 percent write prompts and instructions more often, 38 percent prepare context for the tools. Only 30 respondents see no significant new tasks. So new work appears alongside the automated kind. AI assists with code, tests, documentation and research, and the control work grows in parallel. Whether the total volume of work rises or falls cannot be derived from this survey. One thing does stand out: across more than 1,200 free-text responses from the 383 participants, not one spontaneously reports job cuts caused by AI. The question was not put directly, which makes this an interesting null finding rather than evidence about the labour market. At the same time, only 4 percent experience their work as devalued, and 45 respondents say they would welcome additional conventional developers in their teams. Development work is also not a fixed pie. When AI lowers the cost of building, applications become viable that would previously have failed on budget. Automation can therefore create new development work instead of merely replacing existing work (see also the [Jevons paradox](https://en.wikipedia.org/wiki/Jevons_paradox)). Anyone whose strategy thinks only in terms of replacement underestimates that dynamic. What matters is identifying where the additional value sits and funding it deliberately. The new work comes at a price. 67 percent describe AI-intensive working days as denser, more demanding, or equally demanding at higher output. Only 6 percent feel noticeably relieved. AI does automate routine, but the capacity it frees is quickly taken up again by review, steering and rising expectations. So far the productivity gain shows up less as relief than as a denser working day. > Development work is turning into review work. Leave it out of the plan and you still get it: unbudgeted, invisible, and paid for out of the substance. ## Skill Shift: From Producer to Accountable Reviewer Putting the two sides next to each other shows most clearly which skills are gaining value and what AI is actually used for: | Gaining value / newly emerging | Share | What AI is used for | Share | | -------------------------------------------------------- | ----- | ------------------------- | ----- | | Reviewing and curating AI output (new task) | 68% | Writing new code | 76% | | Domain and business knowledge | 60% | Research and summaries | 44% | | System design and architecture | 56% | Refactoring and migration | 35% | | Requirements and stakeholder communication | 48% | Documentation | 28% | | Test strategy and [evaluation](/en/glossary/evaluation/) | 30% | Debugging | 26% | The left column shows what is newly emerging or gaining value, the right what respondents use AI for, which ranges from assistance to full handover of a task. The pattern is clear either way: judgement, control and context competence are gaining weight, while AI is deployed mainly on the producing side of the work. Prompting is a new activity for 48 percent, but only 32 percent count it as a skill that gains them value. Context and domain knowledge matter considerably more. That supports the hiring logic of the "domain expert with an appetite for experimentation" I described [after PyCon DE & PyData 2026](/en/blog/evaluation-is-architecture/). The shift looks different depending on the field of work. Data engineers move furthest towards subject-matter depth and business translation. Web developers increasingly take on the architecture and control of generated code. ML engineers use AI most intensively and most often build [evals](/en/glossary/evaluation/), context pipelines and agent control. A possible new role profile is taking shape here: the eval engineer, already visible in practice before the job ads catch up. ## Trust Is Growing Faster Than Control The most revealing finding sits outside the jobs debate. Trust in AI output has recently risen for 58 percent, while the control structures are not keeping pace. 89 percent review AI output manually, but only 30 percent have a formal eval pipeline. 14 percent mostly go on gut feel. Accountability is usually unresolved as well. 71 percent say it is not explicitly defined who answers for a defect caused by AI. In practice it lands with the developer who accepted the code. Only 11 percent have a documented framework with mandatory review, sign-off and audit trail. This is exactly where control debt builds up: production is automated faster than review, accountability and governance can grow. Like technical debt it stays invisible at first, right up to the point where it falls due. The data also shows that AI adoption is not merely a technology problem. Asked for the biggest bottlenecks, respondents name unclear responsibilities (40 percent) plus security, data protection and the EU AI Act (36 percent), not the tools themselves. The decisive problems arise at the interfaces between the executive level, engineering and compliance. The 59 incidents described make that concrete: scope changed silently, hidden code duplicates, proposals to delete production data despite guardrails. One respondent sums up the underlying pattern: > "It happens fairly often when generating code that package versions are simply invented […] If you trust that blindly, you quickly end up in a spiral of hallucinations and debugging in completely the wrong place." (translated from German) The new review work only becomes durable as a system: as a [harness](/en/glossary/harness/) of evals, audit trail and governed release. I described the same pattern through trust boundaries in [Agents in Production](/en/blog/agents-in-production-2026/). Left as individual heroics at a departmental boundary, the control work simply drains away. ## Rethink, Don't Replicate A finding is not yet an instruction, so here is the reading. Map existing processes one-for-one onto AI and you automate their weaknesses along with them. The *Economist* [describes this pattern using India as its example](https://www.economist.com/asia/2026/06/28/why-cant-indias-government-build-a-decent-website/): plenty of parties solve their part of the task, yet without integration no working whole emerges. AI agents carry the same risk. Each individual result can be plausible while accountability and control over the end-to-end process are missing. The answer is not more manual review but control as a system, plus the willingness to reorder processes and accountability. How harness, evaluation and governance fit together is something I set out in [The New AI Architecture in 2026](/en/blog/ai-architecture-2026-harness-evaluation-open-source/). ## For Decision-Makers, This Means ### First: Budget Control Work as Real Work Reviewing, curating and steering AI output is daily reality for 68 percent. That work takes time and therefore belongs in the plan. Leave it out and you risk overload and lost quality. That is the first instalment falling due on the control debt. ### Second: Organise Governance as a Cross-Cutting Function The bottleneck is not only technical, it is unresolved accountability. For 71 percent it is not explicitly defined who answers for defects caused by AI. Durable AI use therefore needs a shared accountability model across the executive level, engineering, security and compliance: mandatory review, sign-off and audit trail, carried jointly instead of delegated downwards. ### Third: Scale Manual Review With Evals 89 percent review AI output manually, but only 30 percent have a formal evaluation. As AI use grows, that review load can become the bottleneck. Manual review stays important, but it needs a scalable, reproducible and evidenced complement: your own versioned eval sets with clear subject-matter ownership. ### Fourth: Rethink Profiles and Collaboration The survey measures no employment effects, but it does show a clear skill shift. Domain knowledge, architectural judgement and eval competence are gaining weight. Roles, skill development and recruiting should be aligned to that. Nor is this a one-way street for developers. AI also enables business functions to take on tasks that previously depended on development teams, and new forms of collaboration follow from that. The decisive question is less which task belongs rigidly to which role, and more this: who does what best, and how does the team work with AI to greatest effect? That several respondents miss having additional conventional developers in their teams underlines the point. Automation does not automatically remove the need for development. The common denominator across these four points is a question of organisation, not of tooling. Control debt accumulates at the seams between the executive level, engineering and compliance, wherever accountability is not clearly defined. That gap does not close from inside a single department. It takes someone who is taken seriously from the executive floor down to the developer team and who translates in both directions. That is my work as an independent advisor: connecting the levels, putting what AI actually delivers into perspective, and turning control work into governance that holds. If you want to know where your own control debt is largest, the [Boardroom-to-Code Session](/en/boardroom-to-code-s26/) is the structured way in. > The more intensively developers use AI, the more important the work that hardly any company has organised: reviewing, assigning accountability, securing quality. --- ### The New AI Architecture in 2026: Harness, Evaluation, Open Source URL: https://hendorf.com/en/blog/ai-architecture-2026-harness-evaluation-open-source/ Published: 2026-05-26 Tags: ai-strategy, enterprise-ai, ai-architecture, mlops, llmops, regulated-industries, takeaways > Build with the three shifts in 2026 and you'll have several use cases in production in 2027. Build without them and you'll still be stuck on your first pilot. The most important movement at PyCon DE & PyData 2026 wasn't in any single talk. It was a shift in tone: away from "look at what's possible", toward "running since February". The industry hasn't arrived. But it has digested its first AI wave and pulled three architectural shifts out of it: from the model to the [harness](/en/glossary/harness/), from test to architecture ([evaluation](/en/glossary/evaluation/)), from budget option to sovereignty lever (open source). 2026 is the dividing line between programmes that build with these shifts and programmes that keep building on 2024 logic; the delivery gap will be plain to see in 2027. ## At a Glance - Three shifts that surfaced across this series come together into one picture: AI architecture in 2026 calls for different thinking than 2024. - From model to [harness](/en/glossary/harness/); from test to architecture ([evaluation](/en/glossary/evaluation/)); from budget option to sovereignty lever (open source). - What the conference did not solve: LLMOps stack standardisation, review capacity, clear skill paths. Saucedo's survey data point: roughly half of organisations still have no productive ML monitoring. - 2026 is the dividing line between programmes that build with these shifts and programmes that build on 2024 logic. The delivery gap will be visible in 2027. ## Shift 1: From Model to Harness In 2024, the central question in many programmes was: which model is best? In 2026 it isn't. [Sebastian Raschka](https://sebastianraschka.com) put it on a single line in the fireside chat *Stop Waiting, Start Shipping*, which I moderated, and the line cuts sharpest for Europe: post-training and [harness](/en/glossary/harness/), not a base-model race. The case in point was Cursor's Composer 2: running in production, materially better than most coding LLMs on the market, yet not the product of proprietary pre-training. At its core: an open base model, Kimi 2.5, with additional reinforcement learning on top. Cursor co-founder Aman Sanger confirmed this publicly in [March 2026](https://techcrunch.com/2026/03/22/cursor-admits-its-new-coding-model-was-built-on-top-of-moonshot-ais-kimi/); Moonshot AI described the procedure as "continued pretraining & high-compute RL training" on Kimi 2.5. The production gain didn't come from picking the right model. It came from the layer above it. The same pattern shows in Claude Code: strong as a coding agent because Claude Code is a very good harness for programming work. Under the hood, less model magic than fine-grained division of labour between specialised small programs: code search, patch generation, test execution, diff generation, sandbox invocation. The model orchestrates; the tools do the work. For agent architectures, the shift cuts twice. The programmes running production agents in 2026 have strict context management, deterministic fallback paths, and evaluation-coupled releases. They treat trust boundaries as contracts, not slogans. They have (as [Gabriela Bogk](https://www.linkedin.com/in/gabriela-bogk/) crystallised in her keynote) asked "what is the blast radius?" of every autonomous step before signing it off. Anyone still benchmarking models in 2026 is optimising at the wrong layer. The lever sits two layers higher. ## Shift 2: From Test to Architecture (Evaluation) The second shift is unshowy but central: teams no longer build first and then check whether it works. They test first and then build with intent against what they want to measure. In 2024, evaluation was a downstream check in many programmes. In 2026, it sets the tempo of productive AI development. [Frank Rust](https://www.linkedin.com/in/frankrust/) and [Thomas Prexl](https://www.linkedin.com/in/thomasprexl/) described in their talk *It Works on My Machine* what that looks like in practice: before the first line of code is written, they collect 100 or more real user questions with correct answers and source references. The set is reviewed against the baseline at every sprint review, alongside the power users. [Andrei Beliankou](https://www.linkedin.com/in/andrei-beliankou/) and [Evgeniya Ovchinnikova](https://www.linkedin.com/in/evgeniya-ovchinnikova/) (E.ON) showed the operational layer above: three observability stacks running in parallel, span-level tracing, cost breakdown, and a pointed warning about the failure mode of the evaluation loop that any programme without this discipline will run into. The consequence is different from what most teams assume: the strategic core of a programme isn't the model: it is the eval set. > Models change every quarter. Eval sets stay. Eval sets are the most valuable asset of a 2026 AI programme, and they belong in week one, not in phase four. That moves the hiring category as well. The bottleneck hire in 2026 isn't the ML PhD: it's the domain expert with an appetite for experimentation, a thesis [I anchored to evaluation discipline](/en/blog/evaluation-is-architecture/). Eval discipline lives on domain depth, not ML depth. ## Shift 3: From Budget Option to Sovereignty Lever The third shift redefines the relationship between open-source stacks and enterprise strategy. In 2024, open source was the budget option in many DACH boardrooms: free, risky, somehow not "serious". In 2026, that label no longer sticks. Open source is the architectural form in which data control, auditability and strategic post-training converge. Three axes carry the move. ### Data Sovereignty Through Local Models Locally-run models are in 2026 the more obvious choice for sensitive workloads, as Bogk confirmed from a CISO's perspective. Confidential tickets, code repositories, contract drafts no longer leave the organisation's own control plane. [Sovereignty](/en/glossary/sovereign-ai/) moves from the strategy paper into the stack. ### Audit Sovereignty Through Open Models [Sylvain Corlay](https://www.linkedin.com/in/sylvaincorlay/) ([QuantStack](https://quantstack.net/)/[Jupyter](https://jupyter.org/)) at the Open Source as a Business panel (which I moderated) articulated what hits regulated industry as squarely as it hits science: black-box tools are structurally unfit wherever traceability is mandatory. Model weights in your own hands, audit logs at the inference layer, inspection of model behaviour: none of that works without open models. ### Strategic Sovereignty Through Post-Training [Yann Lechelle](https://www.linkedin.com/in/ylechelle) ([probabl](https://www.probabl.ai/)/[scikit-learn](https://scikit-learn.org/)) on the same panel delivered the economic clarification: open source is not a business model: it is a distribution, community, governance and marketing asset. From that emerges a different question than "open or closed?": which layer do we control, and which do we delegate? Differentiation is built in post-training on proprietary data, on a base model you don't have to fund yourself. The European open-source business ecosystem (probabl, QuantStack, spaCy and others) is available as a partner market for exactly these sovereignty programmes. The detailed view sits in the [sovereignty piece of this series](/en/blog/open-source-sovereignty-2026/). Anyone still framing the 2026 choice as "make-vs-buy between US hyperscaler and own model" is overlooking a substantial market. ### LLMOps Stack Standardisation [Alejandro Saucedo](https://www.linkedin.com/in/axsaucedo/) (Zalando) brought an uncomfortable number from his *State of Production Machine Learning Operations* survey: roughly half of organisations still have no productive ML monitoring. What's missing in the LLMOps stack in 2026 is exactly what slowly took shape in the MLOps stack between 2018 and 2022: shared standards, shared patterns, shared tooling expectations. OpenTelemetry standards for GenAI are a start, but the field is heterogeneous and will stay that way for a while. ### Review Capacity at Ten-Fold Code Growth The [New York Times documented in April 2026](https://www.nytimes.com/2026/04/06/technology/ai-code-overload.html) the case of a financial services firm that jumped with Cursor from 25,000 to 250,000 lines of code per month and built up a review backlog of one million lines. The gap between code generation and review capacity is unresolved in 2026. The conference named it; it did not close it. ### Clear Skill Paths for Domain Experts If the bottleneck hire in 2026 is the domain expert with an appetite for experimentation, those people need a learning path. There is no structured one today. Bootcamps don't hit the profile, and neither do classical ML degrees. What's emerging in 2026, at best, is a mentoring model: external advisory meets internal domain. That only scales so far. ## 2026 Marks a Dividing Line in Corporate AI Strategy. Two Very Different Kinds of Programme Are Emerging The first has understood that the contest is no longer settled by larger models and more compute alone. They invest deliberately in the layers that actually matter for their business: tuning and adapting models to their own data and processes, quality assurance for AI output, integration with existing systems, and an architecture that builds compliance, data security and traceability in from day one. These firms make AI sovereignty practical. They don't just talk about it; they run early applications on controllable in-house or open stacks. They document how their systems are tested, which data may be used, where the risks sit, and how decisions remain auditable. And they don't only look for classical AI researchers; they look above all for people who combine domain knowledge, technical understanding and an appetite for experimentation. ## For Decision-Makers, This Means ### First: Shift Investments from Base-Model Compute to the Harness Programme budget does not belong in the next vendor comparison; it belongs in post-training, harness and the evaluation pipeline. Anyone still optimising on model choice in 2026 is optimising two layers below the lever. Cursor's Composer 2 is the case in point: the production gain came from the layer above the base model, not from picking the right base model. ### Second: Shift Hiring Profiles from the ML Market to Domain Expertise Domain experts with an appetite for experimentation are the bottleneck hire in 2026, not ML PhDs. Eval discipline lives on domain depth, not ML depth, and the talent market for that is available, often less expensive, and a better fit for most programmes. Anyone who accepts that opens up a different talent pool than the competition. ### Third: Make Sovereignty Operational, Not Rhetorical Sovereignty does not belong on [a board slide](/en/boardroom-to-code-s26/); it belongs in the stack: local models for sensitive workloads, open models where audit is mandatory, post-training on proprietary data for strategic differentiation. Several use cases running on your own stack, with a documented evaluation trail and a defensible compliance frame: that is the operational mark of 2026. > 2026 is no longer about who has the biggest model. What decides is who takes harness, evaluation and open-source sovereignty seriously as architecture. --- ### Mastery, not Ownership: Where AI Sovereignty Actually Takes Shape in 2026 URL: https://hendorf.com/en/blog/open-source-sovereignty-2026/ Published: 2026-05-18 Tags: open-source, ai-strategy, digital-sovereignty, enterprise-ai, eu-ai-act, regulated-industries > Sovereignty doesn't belong in a strategy paper. It belongs in the stack. In 2026, AI sovereignty emerges not from ownership but from mastery of critical stack layers. Data, inference, post-training, evaluation, compliance/audit, and operations are the six axes that decide how sovereign an organisation actually operates in 2027. The European reflex of 2024 ("We need our own ChatGPT") framed sovereignty as a question of ownership. By 2026, the question has matured: which layer do we keep, which do we delegate deliberately? Without that differentiation, sovereignty stays in the strategy paper, not in the stack. ## At a Glance - Sovereignty in 2026 is a question of mastery, not of ownership: which layers of the AI stack a company controls itself determines how sovereign it really is in practice. - Six stack layers are the decision axes in 2026: data/residency, inference, post-training, evaluation, compliance/audit, and operations. Each can be kept in-house, delegated, or run hybrid, but not ignored. - The boardroom dichotomy, "proprietary = enterprise-ready, open source = hobby project", no longer holds in 2026. [Open-source software dominates the cloud](https://dataintelo.com/report/linux-operating-system-market), often wrapped as SaaS. Open-weight stacks run in production in regulated industries, given a correctly sorted layer architecture. - What stalls sovereignty is rarely the software or the model: procurement reflexes on "vendor with SLA", skill gaps or the outsourcing of responsibility, the persistent security myth, the unanswered day-one plan. ## Open Source Is Not a Business Model: It's a Question of Control [Yann Lechelle](https://www.linkedin.com/in/ylechelle), Executive President & Chairman of [probabl](https://www.probabl.ai/) and former CEO of [Scaleway](https://www.scaleway.com/en/), put the decisive sentence on record: > *"Open source is not a business model. It's a distribution, it's a community, it's a federation, it's a governance concept, it's a marketing asset."* In decision-making language: you are not buying free software. You are deciding where control and adaptability matter for your business, and where an external partner runs the layer more cheaply and cleanly than you can. The question is no longer "open or closed?": it is which layer of the stack do we keep, and which do we delegate deliberately? Ines Montani ([spaCy](https://spacy.io/)) put the related boardroom myth to rest: > *"People use open source not because it's free, I think that's one of these misconceptions, but because it's flexible, it's extensible, you can program with it. Composable."* Companies use open stacks in 2026 not to save money but because it is the only way to retain control of the layers above the base model. The cost calculation no longer backs the old picture either: are the total costs of training, support, and customisation for proprietary software really lower than investing in your own skills? Rarely, and the comparison overlooks what an open stack adds on the other side of the ledger: control over data, inference, eval, and roadmap. A consequence of this: running open-weight models on an open-source stack in 2026 isn't ideology: it is architectural logic. Pick open weights, then layer proprietary orchestration underneath, and you hand back part of the control you just gained: the audit chain breaks where inspectability ends. Europe as an ecosystem is often invisible: the maintainers of critical open-source libraries sit in the EU. Three OSS titans without which no enterprise stack runs, scikit-learn, spaCy, Jupyter, come straight from this ecosystem: probabl carries scikit-learn, the spaCy team works out of Berlin, QuantStack maintains Jupyter. By 2026 these actors plus Mistral and many others are a credible service market too. Not a replacement for US Big Tech, but a different infrastructure architecture, one already in production. Six layers make the question concrete. ## The Six Stack Layers: Keep or Delegate? ### 1. Data / Residency Which data may leave your control perimeter? Local models for sensitive workloads (confidential tickets, code repositories, contract drafts) stay on your own infrastructure. External APIs for non-sensitive paths reduce operating overhead without giving up the core. The maturity of the local stack in 2026 covers internal knowledge bases, coding assistance on confidential code, classification in regulated pipelines, document-grounded Q&A, not frontier-grade, but sufficient, and without data flowing to US servers. ### 2. Inference Where does the model request run, and who sees the logs? Your own inference infrastructure for sensitive paths is practically achievable in 2026 with locally runnable [open-weight](/en/glossary/open-weights-vs-open-source/) models. Runtime audit logs stay in-house. With an external API, this layer stays with the vendor, even if it is now contractually fenced. Hybrid operation is the rule: local for sensitive, API for scalable. ### 3. Post-Training Where does differentiation come from? This layer is hard to delegate, because the business itself lives here. Cursor's Composer 2 is the case study, covered in depth in the [*Stop Waiting, Start Shipping* piece](/en/blog/stop-waiting-start-shipping/) in this series: at its core Kimi 2.5 with additional reinforcement learning on top, publicly confirmed by Cursor co-founder Aman Sanger in March 2026. Composer 2 illustrates the pragmatic approach in miniature: building value on the shoulders of giants beats every reflex to roll your own base model. Value is created not by the choice of base model, but by proprietary data, evals, [harness](/en/glossary/harness/), and product discipline. Ines Montani put the economic logic behind it on record: > *"It's like software 2.0. You have code and data, and a lot of the value is in the data. So we can provide the code basically for free and open source and then focus our product offering on the other part that's more custom."* The base model is the shared code layer. Your own data and the post-training on top of it are the layer where differentiation arises. Anyone delegating it is delegating their business. ### 4. Evaluation How do you measure model behaviour, in-house? [Eval](/en/glossary/evaluation/) discipline, your own eval suite, and a [harness](/en/glossary/harness/) under your own control are the layers in 2026 where domain-specific quality becomes visible and a wrong model choice surfaces early. Without eval discipline, every sovereignty claim stays an assertion and every model migration a gut call. Cheap to build, expensive to neglect. ### 5. Compliance / Audit Can we inspect the model when the regulator asks? [Sylvain Corlay](https://www.linkedin.com/in/sylvaincorlay/) ([QuantStack](https://quantstack.net/), [Jupyter](https://jupyter.org/)) put the argument in its most fundamental form: > *"Anyone who's done science at some point in their life understands that there would be a huge contradiction in trying to understand the world like physics or biology or anything with a tool that you don't have the right to understand or look into."* > > *"The core of the logic and of the intelligence will have to be open source. Otherwise, nobody is going to buy it. And that's a big part of the trust that people have in the software. It's that it's been audited by experts."* Banking supervisors, insurance regulators, and internal audit functions follow the same logic. Corlay's second sentence is the commercial consequence of the first, and the stronger argument in the boardroom. What open weights deliver: model weights, your own inference, runtime audit logs, control over data flows. What doesn't come automatically: full training-data auditability. With Qwen, DeepSeek, Kimi as with Llama, Gemma, Mistral, training-data provenance is rarely fully documented. Anyone building a compliance argument on provenance has to make the open-weight / open-data separation explicit. Proprietary enterprise providers are no longer pure black boxes: OpenAI Enterprise, Anthropic Claude for Work, and Azure OpenAI offer data residency, zero-retention, BAAs, and inference-level audit logs. What remains: no model weights in your hands, no training-data control, no direct access to model behaviour. Sufficient for many use cases, not for model inspection or in-house post-training as part of the compliance argument. ### 6. Operations What do we do on day one if a central provider falls away? QuantStack has answered the question operationally: migration of all services to European providers, a "contingency forge" as a GitHub mirror so the team can keep working the next day if the main access drops. Sylvain Corlay calls this *economic survivalism*. > *"We've migrated all of our services to European providers … we even have a contingency forge with pretty much everything we touch on GitHub mirrored so that we have something to do the next day, if we are shut down, basically."* Translated to enterprise stacks: service partners instead of vendor licences, documented backup paths, a day-one answer that sits in the boardroom, before it is needed. That capability alone is leverage at the negotiating table. This layer decides whether the architecture still holds in 2027 or breaks at a single vendor decision. ## An Honest Reading of Model Choice The model choice matters, but it is not what matters most. What matters is control of the layers above it: data, eval, harness, post-training. More on this in the [*Stop Waiting, Start Shipping* piece](/en/blog/stop-waiting-start-shipping/) in this series. ## The Organisational Blockers What stalls sovereignty is rarely the model: it is organisational blockers. Four are the most stubborn in 2026. ### Procurement Reflex on "Vendor with SLA" Procurement teams are conditioned to "vendor with SLA". An open stack needs a different contract model: service provider, engineering advisory, internal responsibility. Programmes that fail to reframe procurement early do not fail at the model: they fail at the contract template. Ines Montani delivered a warning that belongs in every procurement discussion: > *"You adopt a project and then the startup behind the project goes and raises a lot of money, and they pivot here, they pivot there, and in the end the open source is kind of collateral."* VC funding is not proof of stability. Stability sits in governance. My own point on this: open-source licences are typically irrevocable. After a pivot or insolvency on the vendor side, the right to use, fork, and develop the software remains, a structural safeguard no vendor contract delivers. ### Skill Gap: Operations, Training, OSS Composition The skill bottleneck in 2026 rarely sits with the ML PhD. It sits in three mutually dependent disciplines: operations (local inference stacks, audit logs, documented migration paths, the day-one plan), training (post-training, continued pre-training on your own corpus, eval suite, harness discipline), and OSS composition, the often overlooked key discipline of selecting, integrating, and maintaining the open components into a coherent architecture. Without these three disciplines, in-house or at a service partner, no layer programme runs. The thesis "[domain expert with appetite for experimentation before ML PhD](/en/blog/evaluation-is-architecture/)" still holds, but does not suffice on its own. Anyone delivering sovereignty in earnest needs a bridge between internal engineering substance and external advisory. ### The Open-Source Security Myth The notion that open source is less secure than proprietary is empirically hard to support. The code is open, and therefore auditable. Proprietary code typically isn't: what happens at the vendor stays at the vendor. On the security question, the point goes to open source: more eyes watch the code than at any single vendor, CVE processes are transparent, the patch path is in the open. Security doesn't end there. It also depends on technically enforceable guardrails: tool permissions, filesystem access, API keys. With an open stack, you can inspect that layer yourself. With a black-box API, you depend on trust in the vendor. ### The Unanswered Day-One Plan What do we do on day one if AWS, Azure, or a central service provider suspends our accounts? Most enterprise stacks have no rehearsed answer in 2026, not from technical impossibility, but from the assumption it won't happen to us. Sovereignty that lets this assumption stand is theatre, not architecture. ## For Decision-Makers, This Means ### First: Take a Sovereignty Inventory Along the Layers Which workloads must run locally, which may go into the cloud? Which layer do we keep, which do we delegate to a service partner? This inventory is the foundation of any sovereignty architecture, and in 2026 a concrete working document, not a strategy paper. How that decision gets made in a structured way is exactly what a [Boardroom-to-Code Session](/en/boardroom-to-code-s26/) is for. ### Second: Reframe Procurement Don't search for a "vendor licence", search for a "service partner for the open stack". A different contract logic, more shared engineering responsibility, and an explicit examination of the partner's governance, not just its SLA logo. ### Third: Build a Use-Case Island, Not a Pilot A "pilot" is defensive and rarely becomes architecture. A use-case island is an internal programme with a local model, your own post-training, your own eval discipline, and your own harness, on a clearly bounded use case. The island is the smallest complete implementation of the layer map, and therefore the shortest path from strategy paper to stack. ### Fourth: Prepare a Day-One Answer Before You Need It What do we do if a central provider falls away tomorrow? A documented migration path, a mirror of the critical repositories, a service partner that can step in, in 2026 not paranoia, but architectural duty. Corlay's *economic survivalism* belongs in every board paper that claims sovereignty. --- Sovereignty in 2026 is no longer a question of ownership. It is a question of mastery. Which layer of the stack a company controls, and which it deliberately delegates, determines how sovereign it really is in 2027. --- ### Regulated, Interconnected, Stalled: What's Blocking AI Projects in Five Industries URL: https://hendorf.com/en/blog/problem-clinic-pyconde-2026/ Published: 2026-05-11 Tags: regulated-industries, ai-strategy, governance, communication, open-source, sbom, model-validation, enterprise-ai > In regulated industries in 2026, AI rarely fails on missing technology. The harder blockers sit upstream: a missing shared language between engineering and compliance, a missing binding standard for "good enough", and a missing mandate to settle these questions across functions. 20 practitioners from five regulated industries - banking, pharma, medical-product development, healthcare IT providers, critical infrastructure - spoke for one hour under Chatham House Rule. Not for the stage, not for a recording, not for slides. About the points where their programmes actually get stuck. That confidentiality was decisive: without it, many statements would have stayed in the usual conference register. Instead, this wasn't about curated success stories but about the operational bottlenecks behind the programmes. The core observation was unambiguous: the industries differ in their rulebooks, but not in their pattern of blockers. Across all five themes, the dominant blocker was almost never the tooling itself, but a combination of three organisational bottlenecks - a missing shared language between engineering and compliance, a missing binding standard for "good enough", and a missing mandate to decide across functions. ## At a Glance - Three bottlenecks dominated across industries: **language** (engineering and compliance talk past each other), **threshold** (nobody operationally names what "good enough" looks like) and **mandate** (no role has the standing to negotiate between the functions). - Five themes recurred in almost every industry: SBOM and sub-dependencies, model validation, the perception of open source, engineering-to-QA communication, AI-generated code in review. - The unresolved question nobody in the room had cracked: LLM validation in productive, life-adjacent systems under regulatory oversight. - The practical implication: solve language, threshold and mandate in one industry and you've already half-built the blueprint for the other four. ## Language: Engineering and Compliance Talk Past Each Other In nearly every account, the same fracture point surfaced: engineering does not understand the vocabulary of compliance, compliance does not understand the vocabulary of engineering, and nobody translates. This isn't a soft-skill problem; it's a shared technical language that, in most houses, simply doesn't exist. The result is that requirements get formulated past each other, acceptance criteria stay vague, and each side experiences the other as a brake rather than a partner. A second observation in the same vein, also from the room: hardly any tech conference has a *communication track*. There are tracks for architecture, for ML, for security, for cloud, for DevOps. But not for the one discipline at which most regulated programmes actually fail in day-to-day work. What that means in practice: as long as a regulated AI programme has no shared language between engineering and compliance, every architectural debate is a proxy fight - and every delay costs more than the underlying problem. ## Threshold: "What Would Actually Be Good Enough?" One voice in the room put it as a question that has not left my head since: *"What would actually be good enough?"* In many corporates, that question has no clear counterpart. When compliance can't name the threshold operationally and engineering isn't allowed to ask for it, every architectural decision becomes a negotiation without a scale. The consequence isn't a too-high or too-low bar; it's arbitrariness - what gets accepted today is insufficient tomorrow, and nobody can say why. The problem rarely lies in nobody wanting quality. It lies in the fact that quality doesn't get translated into decidable criteria. "Safe", "robust", "auditable" or "compliant" are valid as goals, but for a development team they're not sufficient. A team needs the answer to the operational question: how do we know that this solution is good enough for this context? What that means in practice: without an operationally formulated "good enough" - measured, shared, documented - even the best stack can't be cleared for release. The bottleneck sits before the code, not inside it. ## Mandate: Who Has Standing to Negotiate Between the Functions? The same anchor came up several times in the room: innovation teams as an institutionalised bridgehead inside regulated corporates. They are neither pure engineering nor pure strategy. What matters is not their name, but their mandate: they are allowed to translate, negotiate and escalate between engineering, compliance, QA, security and senior leadership with binding force. Where these structures exist, new stacks make it through conservative IT setups. Where they don't, Python stays in the sandbox - for organisational reasons, not technical ones. Then every team can do good local work and still fail at the handover into productive responsibility, because nobody is allowed to carry the decision across functional boundaries. What that means in practice: a mandated bridge function is the precondition for language and threshold to be negotiated at all. Without it, both bottlenecks remain structurally unresolvable. ## SBOMs and Sub-Dependencies - Who Owns the Supply Chain? An SBOM (Software Bill of Materials) is the ingredients list of a software product: every library, every module, every third-party component that ends up in the running system. Sub-dependencies are the ingredients of those ingredients - software is built on software, often five or ten layers deep. Anyone signing off in a regulated industry that a system is secure and auditable has to know not only what is directly included, but also what those included components themselves bring along. A vulnerability three layers down is not a footnote; it falls under the same accountability. The discussion in the room was not about tools - those exist - but about accountability and supply chain. Who maintains the SBOM? Who escalates a critical vulnerability in a sub-sub-dependency? Who carries the risk when a supplier quietly abandons a component, an acquisition rewrites the roadmap, or licence terms tighten mid-lifecycle? What that means in practice: as long as SBOM stewardship is treated as a tooling question rather than an accountability and escalation question, audit safety is only formal, not operational. ## Open Source - The Pharma Question One of the most striking moments of the discussion came out of the pharma context: the perception that "free" means somebody is stealing on our behalf. It would be tempting to dismiss this as an isolated anecdote. It is, in fact, a widely held picture - and in regulated corporates a consequential one, because it frames open source as a deficit rather than as a structural advantage. Two points are systematically overlooked. First, open-source licences like MIT, Apache 2.0 or BSD are irrevocable. A version released under that licence today is still under that licence tomorrow - no supplier can take it back, no acquisition can change it, no "strategic pivot" can void it. With proprietary components, exactly that is a standard risk. Second, open source cannot raise its price. With proprietary stacks, the next licence round, the next acquisition, the next "repricing" is a fixed part of the TCO reality - with open software, that lever simply does not exist. Auditability, [sovereignty](/en/glossary/sovereign-ai/), vendor independence, total cost of ownership - on every one of these axes, open software is the more sustainable foundation in regulated corporates, not the budget version. The longer argument is carried by the open-stack piece in this series, *[Stop Waiting, Start Shipping](/en/blog/stop-waiting-start-shipping/)*. What that means in practice: in 2026, open source is, for the overwhelming majority of enterprise workloads, the regulatorily and economically more grown-up choice. Setting this picture straight in regulated industries corrects a misconception that produces real costs in investment and compliance decisions. ## LLM Validation - The Unresolved Question LLM validation, at its core, means demonstrating that a language model inside a productive, regulated process is reliable enough to carry the accountability that any other causal piece of software in that context would carry. Classical validation expects causal behaviour, deterministic answers, reproducible tests. An LLM delivers none of those in the strict sense - and that is precisely where [evaluation](/en/glossary/evaluation/) begins as an operational discipline. There was a notable silence in the room. To the question of whether anyone is running an LLM as a core component of a productive system under GxP or comparable validation, no serious confirmation came back. Selective LLM use, yes - document extraction, pre-classification, helper functions. As a core component of a life-adjacent or financially critical process, validated by the standard that would apply to any other causal piece of software: no. That is not a failure of the practitioners. It is the regulatory gap. The transitional question is whether existing risk frameworks from the pharma world - hit-rate models, statistical acceptance under oversight, continuous post-market surveillance - can be transferred to software components. There are precedents on that side. They have not, in 2026, been systematically translated to the software side. What that means in practice: anyone deploying an LLM in productive responsibility needs an eval pipeline that measures daily, not one that gets reached for once at release. The shift from validation as a phase to [evaluation as architecture](/en/blog/evaluation-is-architecture/) is the only bridge that holds at the moment. ## For Decision-Makers ### First: Mandate the Bridge Function My recommendation: build a [mandated bridge function](/en/boardroom-to-code-s26/) before the next AI programme begins - an innovation unit with its own mandate, its own budget, direct access to senior leadership. Give it three explicit rights: escalating goal conflicts, documenting release criteria, and preparing decisions all the way to senior leadership. ### Second: "Maybe" Is Not an Answer My recommendation: make "what would be good enough?" a standard instrument in every compliance escalation, in two stages - which standard formally applies today, and how it would be operationally measured for this specific system. "Maybe", "it depends" or "we're still reviewing it" are not answers to this question; they are the signal that the requirement is not yet decision-ready. ### Third: Translation as a Hiring and Training Discipline My recommendation: treat translation between technical and regulatory language as a discipline of its own in hiring and development plans - not as a soft skill, but as an operational precondition for delivery. Concretely: role profiles that name translation competence; a shared glossary for the terms decisions hang on; and, in every larger programme, a named individual who owns this translation. --- ### Evaluation beats architecture: the discipline that decides whether AI ships URL: https://hendorf.com/en/blog/evaluation-is-architecture/ Published: 2026-05-07 Tags: ai-evaluation, ai-strategy, enterprise-ai, governance, regulated-industries, mlops, llms > The most important architectural decision for LLM systems in 2026 is not the model. It is the evaluation. Productive teams in 2026 don't win because they deploy the most spectacular model. They win because they measure earlier, more continuously, and more domain-specifically. That was one of the clearest lines running through the talks at PyCon DE & PyData 2026: where eval discipline was thought through from the start, the systems hold today. Where it was defined as a downstream phase, teams are often still on their first version. Evaluation, in other words, is not a testing discipline: it is an architectural decision, and the one that shapes every other. ## At a Glance - The most important architectural decision for LLM systems in 2026 is not the model: it is the [evaluation](/en/glossary/evaluation/). - Three shifts that became visible across the PyCon DE & PyData 2026 talks: continuous instead of one-off evaluation, domain-specific instead of generic benchmarks, eval-coupled governance instead of a separate compliance layer. - The right order when building an LLM solution: baseline in context first, then retrieval-augmented generation, then [fine-tuning](/en/glossary/fine-tuning/), then knowledge graphs if at all, not the other way around. - The 2026 hiring profile: domain expert with appetite for experimentation, not ML PhD with an architecture plan. ## Three Shifts That Became Visible in 2026 ### Continuous, Not One-Off The view that evaluation is a step at the end ("we have the model, now let's evaluate it") doesn't hold up in 2026. In the talks describing serious production work, evaluation was a continuous pipeline: with every model swap, every prompt change, every new dataset. The effort only pays off once it is already built in. Programmes that don't do this from the start mostly fail to retrofit it later. ### Domain-Specific, Not Generic Public benchmarks like MMLU or HumanEval have their value for model comparisons buried in the small print. For the question of whether a model holds up in a specific application, they are nearly useless. [Frank Rust](https://www.linkedin.com/in/frankrust/) and [Thomas Prexl](https://www.linkedin.com/in/thomasprexl/) showed in their talk how productive LLM teams approach this: they don't start with architecture, they start with 100 real user questions, correct answers and reliable sources. That set is what improvement is measured against. [Cheuk Ting Ho](https://www.linkedin.com/in/cheukting-ho/) drilled through the same logic methodically in the eval workshop *Do you know how well your model is doing?* own tasks, own metrics, with LightEval as the tool. These internal eval sets are often the most valuable asset of an AI programme in 2026, they outlast every model change and define what "good" actually means. ### Eval-Coupled Governance The old architecture separated two layers: model evaluation on one side, compliance and governance on the other. The productive programmes of 2026 have collapsed that separation. Audit logs, drift detection, regulatory requirements, bias metrics: all run inside the same eval [harness](/en/glossary/harness/). [Andrei Beliankou](https://www.linkedin.com/in/andrei-beliankou/) and [Evgeniya Ovchinnikova](https://www.linkedin.com/in/evgeniya-ovchinnikova/) (E.ON) showed concretely what this looks like: three observability stacks in parallel (Langfuse, Opik, MLflow), tracing of individual spans, cost breakdown along prompt paths, and a typical production failure that any programme without this discipline runs into: competing eval checks (hallucination versus faithfulness) escalate against each other until the retry limit fires, and afterwards no one can say why the answer got worse. [Alejandro Saucedo](https://www.linkedin.com/in/axsaucedo/) (Zalando) put this into a wider arc: roughly half of organisations, according to the current Survey on the State of Production Machine Learning Operations, still have no productive ML monitoring in place. His takeaway after two decades of production ML experience: less hype around models, more robust operational, monitoring and governance structures. Anyone setting up a programme in 2026 that treats this discipline as ornamentation is repeating the mistakes of 2018 to 2022, with larger models and higher costs. In regulated industries this is not comfort but the precondition for a system to be released at all. ## The Hierarchy of Solutions, and Why Most Get Complex Too Early [Sebastian Raschka](https://sebastianraschka.com) described in the fireside chat at the conference a sequence that gets walked the wrong way round in many programmes: anyone building an LLM-based solution should test the simplest form first: load the relevant information into the context, no further complexity, and see how far that gets you. That is the baseline. Only when it doesn't suffice is retrieval-augmented generation worth it. Only when that doesn't suffice either is fine-tuning worth it. Knowledge graphs come, if at all, last. Most teams reverse this. They jump straight into fine-tuning because it is technically interesting. Or they build knowledge graphs because that looks like control. Or they discuss MCP integrations without ever checking whether a simple CLI access solves the same task. This is not expensive because of the extra work. It is expensive because eval discipline blurs: anyone who starts with fine-tuning can no longer say cleanly whether the improvement came from fine-tuning or from something context management would have delivered just as well. The baseline is missing, and with it the basis on which every further investment can be assessed. ## Hiring Profile: Domain Expert Willing to Experiment When a conference attendee asked who a "manager with budget and a sovereignty mandate" should hire, Raschka answered in a way that gets lost in DACH reality: not primarily ML specialists. You need people who have worked with these systems enough to develop intuition, and who are willing to try things out. > "A domain expert who is willing to delegate the boring stuff." That is a different hiring logic from "senior ML engineer with five years of PyTorch". It shifts the focus to two qualities: deep understanding of the business process, and the willingness to pick up a tool that you can learn to operate well enough in two or three months. Anyone who accepts that the decisive ingredient is not ML depth but domain-anchored eval discipline opens up a different talent market, one that fits most mid-market and corporate programmes considerably better. ## Eval Sets as a Strategic Asset Models change, every quarter. Which means the strategic core of a programme is not the current model but the eval set: 200 to 2,000 real examples from your own business process, with answers a domain expert has signed off as correct. With this set, every new model can be measured against your own reality within hours. Without it, you depend on public benchmarks that are rarely relevant to your own question. > Measure, don’t guess requires building your own versioned eval set; otherwise, every model decision remains a matter of belief. A robust composition stratifies by characteristic: typical standard cases, borderline cases (ambiguous inputs), historical difficulties (cases where domain staff themselves disagreed), deliberate stress tests (manipulated inputs, edge cases of the compliance requirement). 200 to 500 cases, roughly evenly distributed, cover most of the reality, the exact number follows the domain, not a rule of thumb. Quarterly maintenance; annotation stays a human responsibility. From accompanying several AI programmes over the past year: programmes without their own eval set spend the first three months optimising things that can't be measured, and then realise they have no reference point for saying when it has got better. The build commits domain-expert capacity and needs supplementary engineering for versioning and pipeline integration. Even so, it is by some distance the most profitable investment in the first phase. ## For Decision-Makers, This Means ### First: No AI programme without your own eval set Eval setup belongs in week one, owned by the business owner, not parked in the ML engineering roadmap. ### Second: No fine-tuning without baseline evidence Jumps into more complex layers need eval evidence, not enthusiasm: baseline, [RAG](/en/glossary/rag/), fine-tuning, knowledge graph, in that order. ### Third: No governance concept without a technical audit trail Compliance and eval belong in the same harness; without end-to-end observability, sign-off in regulated industries is not defensible. How that governance decision gets made in a structured way is exactly what a [Boardroom-to-Code](/en/boardroom-to-code-s26/) engagement is for. ### Fourth: No hiring solely along classical ML profiles Domain expert plus appetite for experimentation is an available, often cheaper, and for most use cases better-fitting profile. Raschka's closing advice was drastic in its simplicity: try things out instead of planning too long. The talk title *Stop Waiting, Start Shipping* was the programme. Anyone starting in 2026 without eval discipline will not be buying themselves a shortcut in 2027. > Stop Waiting, Start Shipping! --- ### Agents in Production: Why Evaluation Matters More Than Model Choice URL: https://hendorf.com/en/blog/agents-in-production-2026/ Published: 2026-05-04 Tags: agentic-ai, enterprise-ai, ai-strategy, governance, regulated-industries, ai-architecture ## At a Glance - At PyCon DE & PyData 2026, teams are showing agent systems that have been running stably in production for three to six months, not just demos any more. - Three patterns connect those systems: strict context management, deterministic fallback paths, [evaluation](/en/glossary/evaluation/)-coupled releases. - Three patterns reliably break: open-ended task specifications, free tool choice, test generation without a domain anchor. - The strategic consequence: in 2026/27, agent architectures should be measured by their [harness](/en/glossary/harness/), not by their model. Anyone putting the next euro into the base model is investing where the lever isn't. ## From Demo to Production: A Quiet but Hard Cut Something at PyCon DE & PyData 2026 was hard to capture in a headline: the tone of the agent talks has turned. Where in 2024 and 2025 almost every demo still opened with "look at what's possible", this year teams talked about "running since February", "had to be re-tuned three times", "fails in exactly these categories". That sounds unspectacular. It is the most important shift of the year. What surfaced there was not new agent euphoria but its opposite: sobriety as a sign of maturity. The teams that ship don't talk about model magic any more. They talk about context budgets, fallbacks, tool boundaries and evaluation. In short: about the harness. By harness here I mean the whole of context control, tool boundaries, eval pipeline, fallback logic and approval model: the operational layer that makes a model production-ready. It is at the same time technical architecture, governance frame and investment decision. Three patterns from the conference recurred in the systems that hold, and three in the ones that dazzle in the demo and reliably break in production. ## Three Patterns That Hold in 2026 ### Strict Context Management Instead of Open Memory The systems that work treat context not as unlimited memory but as a scarce resource with explicit economics. They know what gets in, what stays out, and when it gets pruned. [Sebastian Raschka](https://sebastianraschka.com) put it dryly in the fireside chat: the secret behind working coding agents isn't the model: it's prompt and cache management. The repo history, the conversation, the plan: all of it has to be fed in, but not all at once. Without active curation, you build a system whose behaviour drifts from session to session. Context management, then, isn't a prompting technique. It is state management. ### Deterministic Fallback Paths Every robust system has a path that works without an LLM. That isn't nostalgia for the past, it is the honest acknowledgement that a language model does not get more available, cheaper or more explainable the deeper you bury it in the stack. In regulated contexts the point is non-negotiable: without a deterministic fallback there is no audit trail, no traceability, no sign-off. The fallback isn't the system's emergency exit. It is the proof that the system has been understood. ### Evaluation-Coupled Releases The production-stable teams don't release "when it looks good", they release when the eval pipeline is green. That presupposes the pipeline exists, which in many programmes happens late, often too late. Where eval discipline was built in from the start, the conference showed systems with clear version states and traceable improvement curves. Where it was missing, you saw teams that could no longer say with any precision when their system had actually got better. Eval is not quality assurance at the end. It is the only layer in which "better" has any meaning. ## Three Patterns That Break in Production ### Open-Ended Task Specifications "Write tests for this module" produces tests. They are rarely good. Raschka put it plainly in the chat: agents have "no agency of their own". They respond precisely to precise instructions. Where open-ended tasks are set, you get shallow solutions, which look impressive in the demo because they do *something*, and fail in production because "something" is not enough. [Alina Dallmann](https://www.linkedin.com/in/alina-dallmann/) dissected this precisely in her talk *Beyond Vibe-Coding: A Practitioner's Guide to Spec-Driven Development*: three recurring failure modes appear when you give the AI open-ended tasks: fragmented design decisions scattered across multiple chat sessions; prompt drift, where the conversation develops a life of its own; and hidden assumptions the model makes because no one stated them. Her conclusion is the same one that holds here as an architectural claim: the specification belongs before the code, not inside it. Task specification becomes architecture. Anyone who doesn't see that has a problem that isn't a model problem. ### Free Tool Choice When an agent gets to pick from an open toolbox, behaviour in practice tips into the unpredictable, mostly elegant, occasionally catastrophic. Harald Nezbeda's talk *Building Secure Environments for CLI Code Agents* delivered concrete incidents from practice (more in the "Trust Boundary" section). For non-critical applications, that spread is fine. For any regulated context, any critical pipeline, any automated operation against real data, it is an architecture that gets expensive sooner or later. The systems that hold restrict tool choice drastically, and they check every tool against a clear use-case contract. ### Test Generation Without a Domain Anchor The New York Times documented the case in April 2026: a financial services firm jumped from 25,000 to 250,000 lines of code per month with the AI coding tool Cursor. Within a short period, a review backlog of one million lines built up. Joni Klippert, co-founder and CEO of StackHawk (a security start-up working with the firm): "The sheer amount of code being delivered, and the increase in vulnerabilities, is something they can't keep up with." The consequence: senior software engineers in urgent demand to do the reviewing, and pressure cascading into sales, marketing and support, who have to keep pace with the tempo. Tests that agents write are often shallow. Reviews that humans do don't scale by a factor of ten. That gap will not close on its own in 2026. ## What This Means for Architecture Decisions If the harness is the decisive layer, three concrete consequences follow for programmes starting in 2026: ### Model choice becomes secondary Cursor's Composer-3 (running productively) is one example of why: the production gain came from post-training on an open base model, not from the model choice. Accept this logic and your investment shifts from vendor comparison to harness engineering. The full case-study treatment of Composer-3 sits in the [open-stack piece in this series](/en/blog/stop-waiting-start-shipping/). ### Trust boundaries become explicit Where can an agent act autonomously, where only suggest, where only inform? This is not a detail question. It is architecture. In regulated industries it defines the compliance frame. Everywhere else it decides whether the system can scale. ### Review capacity becomes the bottleneck The factor-of-ten jump in code volume happens automatically once agents start writing productively. Reviewers to check it do not appear automatically. Programmes that don't address this in setup build themselves a piece of technical debt that costs more in twelve months than any savings ever return. ## Trust Boundary as Contract, Not Slogan Saying "trust boundaries" is easy; writing them as a contract is the actual work. This is exactly where most programmes stumble in 2025/26, not because the idea is wrong, but because it never gets made operational. Concretely: for every agent step, three modes are worth distinguishing. Read and suggest (human decides), execute with downstream sign-off (four-eyes principle), or complete autonomously (no human in the loop). Which mode applies to which action is not a technical decision but an [architecture and governance](/en/boardroom-to-code-s26/) one. It has to be encoded in the pipeline, not in the prompt template. The programmes delivering productively in 2026 have explicitly assigned these three modes for every agent step. In most cases the autonomous variant is excluded for more than half of the possible actions, and precisely that is what makes the rest defensible. Where this assignment is missing, every action runs implicitly as "autonomous" until an incident forces the discussion. From running architecture reviews I can add: where this mode assignment is explicitly part of the contractual setup of the programme, it becomes operational. Where it is carried along as an annex or as an implicit architectural decision, it collapses at the first stress moment. [Gabriela Bogk](https://www.linkedin.com/in/gabriela-bogk/), CISO at Mobile.de and a long-time member of the Chaos Computer Club, captured this in her keynote *"Honey, I vibe coded some crypto"* with a formula that ought to become a contract clause in every architecture discussion: **blast radius**. The question that has to sit before every autonomous agent step is not "can the agent do this?", but "what is the worst that can happen if it gets it wrong, and can we absorb that?". Her own Claude Code setup runs in a VM with hand-curated API keys, no access to production data, and the code itself backed up in a Git repo. That is not paranoia, it is the translation of trust boundary into operational architecture. Bogk's second point is central for regulated contexts: prompt-based [guardrails](/en/glossary/guardrails/) are soft. *"Everything that's prompt-based in terms of your guardrails is soft and can be worked around and is prone to injection attacks."* Anyone implementing security through system prompts is building on sand. Hard-coded limits on tool permissions, filesystem access and API keys are the only load-bearing layer: the LLM sits on top, not underneath. [Harald Nezbeda](https://www.linkedin.com/in/nezhar/) made the consequences very concrete in his talk *Building Secure Environments for CLI Code Agents*. The risk profile of running a coding agent unsandboxed on a developer's machine falls into what Simon Willison calls the *lethal trifecta*: private data access plus external connectivity plus acting on untrusted context. Documented incidents from real Claude Code use: wiped home directories, a crypto miner installed via a compromised NPM package. His pattern for it: container isolation plus a man-in-the-middle proxy with its own SQLite-based observability. That isn't paranoid. In 2026 it is the minimum configuration whenever a coding agent goes into production or into regulated contexts. Not waiting for that conversation is the most expensive piece of discipline in agent engineering in 2026. ## So What Reading the conference as "confirmation of the agent wave" misreads the picture. In 2026 agent programmes split into two camps: those with harness discipline and those without. The positioning statement "we are now also moving into agentic AI" (whether on board slides, in strategy papers or in the investor deck) is too cheap in 2026. It says something about the external presentation, nothing about the architecture. The question that decides programmes over the next twelve months is more concrete: what is our harness, who builds it, how do we measure that it holds. Not the model choice, not the vendor, not the pilot budget. Take the harness seriously and you get agents that hold. Skip it and you get demos at production cost. --- ### Stop Waiting, Start Shipping: the Open AI Stack Grew Up in 2026 URL: https://hendorf.com/en/blog/stop-waiting-start-shipping/ Published: 2026-04-27 Tags: open-source, ai-strategy, digital-sovereignty, enterprise-ai, post-training, eu-ai-act, llms At PyCon DE & PyData 2026 I had the opportunity to host a fireside chat with Sebastian Raschka, author of *Build a Large Language Model (From Scratch)*, formerly a statistics professor, today one of the few voices that translates between LLM architecture and practitioner reality without watering down either side. What crystallised in the conversation isn't a set of model benchmarks but three strategic lines that matter more for decision-makers than the next release wave. ## There Is No "Winner Takes All", and That's an Architecture Question The mainstream story still runs as a horse race: OpenAI versus Anthropic versus Google versus the Chinese players. Who wins? Wrong question. Raschka puts it plainly: models are increasingly post-trained for their **[harness](/en/glossary/harness/)**: the specific environment in which they run. Cursor (the US coding startup, last valued at around $30 billion) runs *Composer* in its own agent, at its core a post-trained Kimi K 2.5, a Chinese open-source base model. Claude Code, Codex, and any serious coding agent operates inside a purpose-built harness in which the model is post-trained specifically for the tool interface. > "There's no general model that is gonna do well in all the harnesses." Sebastian Raschka That shifts the build-vs-buy logic. "Which model?" is the wrong opening question in most cases. The right one: *what harness are we building, and which model fits inside it?* Picking a model before settling the harness puts the cart before the horse. ## Europe's Allocation Question: Not the Next Base Model The most prominent European AI-strategy debate ("we need our own European base model") is problematic on two counts. First, capable [open-weight](/en/glossary/open-weights-vs-open-source/) base models already exist: DeepSeek v3 and Qwen 3 (both from China), Llama (Meta, USA). Mistral Large 3 (Europe's most prominent model) is at its core a post-trained DeepSeek-v3 architecture. Second, the actual competitive advantage sits one layer up: post-training, tool integration, domain data, and harness design. When a CIO with budget and a twelve-month delivery window asks me where the next euro should go, the answer is clear: **not into pre-training.** It belongs in the layers where domain-specific, regulatory-defensible value is created: the layers that can actually carry data-protection, [on-premises](/en/glossary/on-premises-ai/) and audit requirements. That is the only lever that delivers technological [sovereignty](/en/glossary/sovereign-ai/) and compliance at the same time. This investment logic only works on open weights: post-training on a closed API isn't really an option. Open source is therefore not preference, it's prerequisite. And one of the most underrated reasons even the labs themselves treat it that way is talent. Asked why Google releases Gemma or OpenAI releases GPT-OSS, Sebastian's answer wasn't marketing or generosity: it was hiring: > "If you hire people who never worked on an LLM because there's no LLM you can work on if it's all proprietary, well, you have to train them from scratch." Sebastian Raschka If even the labs treat open weights as a structural prerequisite for their own talent pipeline, the same logic flows downstream: enterprises building on those weights (and hiring people who've worked with them) depend on exactly that precondition. ## Agents on a Short Leash The second wave of coding agents has gone mainstream. Raschka mentioned a friend whose startup codebase was essentially built with Claude Code, work that would previously have cost two years and a thirty-person team. That is the optimistic reading. The other one: in large codebases (Raschka and I discussed the example of New York's financial firms) code volume suddenly grows by a factor of ten. 25,000 lines become 250,000. All of a sudden the senior reviewers required to sign off in a regulated industry are missing. The bottleneck moves from writing to reviewing. > **Experienced engineers and data scientists do not become less important in this world. They become more valuable.** My recommendation is therefore unchanged: **[agents are great, but they need a short leash.](/en/blog/agentic-ai/)** Concretely: do not aim for end-to-end automation; aim for making existing work better. Raschka puts it crisply: > "Making your work better rather than replacing it." Sebastian Raschka The use cases that scale in regulated contexts follow this exact pattern: enrich existing tests with agent suggestions rather than generate them outright; widen code review with a second pair of eyes rather than automate it; make documentation searchable rather than auto-produce reports. ## The Replacement Debate Is the Wrong Debate Public discussion is dominated by the replacement question: which jobs disappear? Which tasks does AI take over? Sebastian's perspective from research turns the argument inside-out. The work PhD students used to spend days on (hyperparameter sweeps, batch-script wrangling, log parsing, plotting results), the grunt work the ML community half-jokingly calls *graduate student descent* (a play on *gradient descent*), now largely automates away. > "The students now actually get to do science instead of doing busy work." Sebastian Raschka That is not replacement. It is release. The creative capacity of researchers (formulating hypotheses, designing experiments, interpreting results) only becomes available once the tedious layer is automated. The same pattern applies in industry: equipping a team with AI agents does not replace heads, it relieves them of the share of work that never required human judgement in the first place. Value migrates upward, not away. ## Practical Consequences Three recommendations I take from the conversation and continue to sharpen in [active engagements](/en/interim-ai-leadership/): ### First: Harness Before Model Move the model-selection discussion to the *end* of the architecture conversation, not the start. Settle the harness first: tool interfaces, data wiring, security boundaries, audit trail. Which model runs inside is the last decision, not the first. ### Second: Invest in Post-Training, Not Pre-Training Proprietary base models would be wasted capital for 99 percent of European companies. Investment in post-training capability, MLOps, data pipelines and harness engineering, by contrast, is where domain value actually compounds. Anyone who takes sovereignty seriously doesn't build the next GPT. They build the layers where domain value is defended. ### Third: Experiment Instead of Wait Raschka's closing advice is also mine: three small experiments this week beat one big plan next quarter. > "If you plan something very thoroughly it is probably irrelevant tomorrow." Sebastian Raschka Waiting is not an option. Unstructured action isn't either. The difference sits in the discipline with which experiments are framed (small models, sandboxed environments, narrowly defined use cases) and in the reflex to fold what you learn straight into the next architectural decision. --- ### Agentic AI & Automation: Three Months of Coding Agents in Operations URL: https://hendorf.com/en/blog/agentic-ai/ Published: 2025-09-15 Tags: ai-strategy, automation, community-management, enterprise-ai, practical-ai, agentic-ai, conference-operations The operational lesson from three months of working with different coding agents is unglamorous but decisive: agents create value when tasks are well documented, narrowly scoped, and verifiable. They create risk when introduced without explicit governance into processes that require judgment, data quality, or multi-step dependencies. The central question is therefore not whether agents can be used productively, but under which governance framework. Identifying suitable use cases and defining the review structure is not a technical detail. It is a strategic decision. ## At a Glance - **Setting**: 1,500 attendees, three months of intensive work with coding agents in the run-up to PyCon DE & PyData operations. Not a controlled experiment, but an anecdotal stocktaking across several models: Claude Code, Gemini, Qwen Coder Plus, Codex. - **Core finding**: Agents perform reliably on tightly defined, well-documented tasks, and break precisely where you'd least expect them to: data normalisation, multi-step pipelines, well-documented security patterns. - **Strategic framing**: Research puts generative AI's share of value creation in analytical programmes at less than 15 percent. The remaining 85 percent stay what they always were: analytics, machine learning, data quality, prediction, governance. - **Risk pattern**: *Augmented Arrogance*: agents deliver confident, formally correct output without ever asking back. The real risk is not the visible [hallucination](/en/glossary/hallucination/), but the unnoticed one inside a productive process. - **Practice**: the *Short Leash Principle*: narrowly scoped use cases, frequent commits, explicit escalation, and a deliberate refusal to chase tool variety. Experienced engineers stay indispensable, as reviewers, not because they write less. ## A Conference as a Field Test PyCon DE & PyData is a fully volunteer-organised conference with 1,500 attendees. In the three months leading up to and during the event we deliberately experimented: which tasks in conference operations can be delegated to AI coding agents: where do the promises hold, where do they break? This setup is a rewarding test bed. A community conference runs on volunteer engagement: domain experts from different fields who come together for a limited stretch, learn from one another and build together, without sitting under enterprise mandates. Audit pressure, change advisory boards and long approval paths fall away; the loop between observation and correction is short. That is also where the actual value of this conference sits: in the engagement of the people involved. Automation therefore lands with a double effect: it takes mechanical work off volunteers and shifts their time toward where programme quality and the attendee experience are actually created. This is not a controlled experiment but an anecdotal account across several models: Claude Code, Gemini, Qwen Coder Plus, Codex. What remains is a subjective but operational stocktaking from real-world ops. With the necessary adjustment to audit and approval frames, it transfers to any organisation thinking today about putting agents into productive processes. ## The Catalogue of Attempts The tasks we delegated to agents read accordingly unspectacular: issuing and tracking speaker and organiser tickets, including cancellations and late additions. Aggregating sponsor and marketing reports. Collecting LinkedIn posts about the conference (more than 250 posts per edition, editorially usable). Auto-generating 120 talk-promotion posts with image, description and link to the video. Producing rough video cuts the day after, using break-slide detection via computer vision. Onboarding 1,500 attendees into Discord with the right roles. Alongside that: individual visualisations, audiobook pipelines, small tools for internal flows. None of this is glamorous. But each of these tasks costs volunteer time when a human does it, and each shows up in equivalent form in any enterprise, only at a different scale and under different compliance pressure. ## Where Agents Hold Up, and Where They Break Unexpectedly After three months of work the pattern is unambiguous, and it depends less on the tool than on the shape of the task. ### Where agents reliably deliver Well-documented APIs (the Pretix ticketing system with its clean REST interface as the standout example: the agent really worked here). Browser automation with defined interaction (LinkedIn data collection, surprisingly good, given that LinkedIn is genuinely hard to scrape). Computer-vision tasks with a clean edge (break-slide detection for automatic video cutting). Boilerplate code. Individual visualisations, without the agent needing deep library knowledge. ### Where they break unexpectedly Precisely where you'd have expected them to succeed. Data normalisation: the question of whether "Acme Corp" and "ACME Corporation" are the same record should be trivial; it isn't. CI/CD pipelines for security scans on Azure, well-documented corpus, should be trivial; it isn't. Multi-step pipelines with heterogeneous tasks, e.g. a four-stage video release process (fetch the video, pull the metadata, publish to YouTube, trigger the follow-up steps), should be trivial; it isn't. Consistent data pipelines with proper error handling in a generic form, also below par. One scene from the three months illustrates the pattern behind these findings. The agent had built a map visualising where conference attendees come from. Visually impressive, technically clean. Only on a closer look at the raw data did it become clear that only about a third of the attendee data was in the dataset at all: the affiliation fields were inconsistently populated, and the agent had not flagged this. What looked like an insight into where the community comes from would have been a visualisation that does not hold up. The lesson from moments like this is not "agents are useless." It is: without a human reading along with experience, any agent in productive responsibility is a liability. A second observation follows from this, and it costs money in practice: data quality remains the load-bearing foundation. A polished visualisation built on a questionable data base now appears faster than ever, and it is more dangerous than an obviously bad dashboard, because it convinces. ## Augmented Arrogance The *Stochastic Parrots* debate (Bender, Gebru and others, 2021) framed the core point, and Yann LeCun has picked up the image: as impressive as large language models look in their agentic packaging, they are powerful pattern generators without understanding. From practice, I see a specific pathology emerging from this, one I called *Augmented Arrogance* in the two underlying talks, *Beyond Agents* and *How We Automate Chaos*. The pattern: agents always produce an answer, and that answer always looks convincing. They do not ask back. They do not say "I am stuck, help me." They produce code that is formally clean and semantically wrong. They decide unprompted that a visualisation would be "better" with three extra elements. They insert steps no one asked for. Even explicit prompt instructions (*"It is fine to ask for help if you get stuck"*) barely move the needle. The behaviour sits deep in the model and the system prompt. This is the opposite of human uncertainty. An experienced junior engineer says: "I am not sure." An agent says: "Here is the solution." Even when it is wrong. In an enterprise context this turns into a concrete risk: a project lead sees highly polished output, trusts the system, and three weeks later a process breaks because a fundamental error has slipped through unnoticed. That is the core risk of unsupervised agent use in companies. It is not the spectacular hallucination. It is the unobtrusive one. ## Short Leash Principle The direct response to this risk is a practice I call the *Short Leash Principle*. It has four components, all distilled from operations. ### First: tight scoping of the use case The agent does not solve "the problem"; it solves a concrete, clearly delimited task. Pull LinkedIn data: yes. "Develop a LinkedIn strategy": no. Produce rough video cuts: yes. Rework the editorial line of the social channels: no. This separation isn't modesty; it is risk management. ### Second: frequent commits, and the right to reset Agents loop, add code without deleting, drift away from the brief. Whoever commits frequently (or asks the agent to do so) can roll back to a clean state instead of disentangling an opaque layering of unrequested features. This is version control as governance. ### Third: tool diet, not tool buffet It is tempting to wire in every available MCP server, every sub-agent, every new framework. In practice it accelerates nothing: it disperses attention and the [context window](/en/glossary/context-window/). Fewer tools, used deliberately, beat the full toolbox. ### Fourth: explicit escalation, and the human as a reader The agent must have a recognisable list of cases in which it is allowed to say: "I cannot do this, a human is needed here." And the human has to read the code, not just write the next prompt. The most honest finding from three months of practice: *agents turn writers into readers.* Whoever prompts more than they read is accumulating debt that comes due later, expensively. > Experienced engineers and data scientists do not become less important in this world; they become more valuable. The strategic consequence: experienced engineers and data scientists do not become less important in this world; they become more valuable. Their centre of gravity shifts from pure building toward the role of reviewer and governance layer. Cutting these roles because "agents will soon do everything" hard-codes a structural weakness into the programme, one that only surfaces in audit or incident, and then expensively. ## The 15-Percent Order of Magnitude This experience meets an estimate I keep encountering in research, and that I see confirmed from work in comparable programmes: in an analytical or software-heavy programme, generative AI covers at most around 15 percent of the value created. The remaining 85 percent stay what they always were: analytics, machine learning, data quality, predictive modelling, security, governance. This sounds like a technical statement. It is in fact an organisational one. It says: anyone investing primarily in agent tooling without continuing to build the classical disciplines is laying an unstable foundation. Agents accelerate a narrow slice. They do not replace the base. Building headcount planning on the opposite assumption, and cutting engineering roles, is a strategic miscalculation that only becomes visible once something breaks. Innovation in this world does not happen inside individual agents. It happens at the seams: between domains, data sources, systems. That requires [people who can read and shape those seams](/en/interim-ai-leadership/). Not people who prompt fast. ## So What The operational lesson from three months of work with various coding agents is unspectacular but load-bearing. Agents create value where tasks are documented, narrowly scoped and reviewable. They create risk where they are placed, without explicit governance, into processes involving judgement, data quality or multi-step dependencies, and they often fail there even when the task looks well-documented on the surface. Anyone who skips this systematic thinking wires risks into their own stack that are hard to undo in audit or in an incident. The question is therefore not whether agents are used productively in an organisation. It is under which governance frame that happens. Both pieces, picking the right use cases and defining the review structure, are strategic decisions, not technical ones. --- ## Selected Slides --- ### Cross-Pollination & AI: Scaling in the Era of Autonomy URL: https://hendorf.com/en/blog/cross-pollination-ai/ Published: 2025-08-15 Tags: Cross-Pollination, AI, Autonomous Systems, Innovation, Knowledge Exchange ## Conclusion: The Future Belongs to the Connected In the era of autonomy it will not be the companies with the best single technology that win, but the ones with the **best networking and the most effective knowledge exchange**. This is not a theoretical ideal. This is what I see live, having worked in the international data-science community for years. I see how a method from space research becomes the standard in finance three years later. I see how an open-source project a junior developer started on a Friday alongside their day job suddenly transforms an industry. I also see the opposite: companies that ignore all these chances because their silos are too deep, their mindset too outdated, their execution backlog too large. Cross-pollination is not a luxury for innovative companies. It is a **survival strategy**. In a world where everything changes exponentially fast, the ability to learn from others and share your own knowledge is the decisive competitive advantage. >The technology will become ever more autonomous. But the human factor of networking will only become more important. We are in a period of transition: **the era of autonomy**. Systems, machines and processes increasingly operate with minimal human intervention. They use AI, robotics and advanced automation to make decisions on their own, adapt to changing conditions and act independently. But how do we scale successfully in this new era? The answer lies in a concept that is older than any technology: **cross-pollination**, interdisciplinary exchange of knowledge. I have been able to watch this force play out for more than a decade, not in a classroom but where it counts: in the international Python and data science community. At conferences like PyCon and PyData I saw astrophysicists talking with fintech founders, industrial-automation engineers adapting algorithms from university researchers, open-source communities solving complex problems without any competitive instinct. That was never theory. That was the proof that real innovation happens at the seams. ## At a Glance **The era of autonomy needs networking.** AI systems get more powerful by the day, but the companies that win are not the ones with the best single technology. They are the ones with the best networking across boundaries. Cross-pollination is not optional. It is the difference between stagnation and leadership. ### Three layers are necessary - **Structure**: create the time and space for exchange. Tech talks, internal wikis, deliberately diverse teams. None of this is free, but it is invested, not spent. - **Culture**: a culture of openness and curiosity cannot be decreed. It is built through continuous accompaniment, through example, through showing that cross-pollination is rewarded. - **Strategy**: all of this needs leadership backing. It needs the clarity that breaking down silos is not just desirable but a question of competitive survival. The ESA-to-fintech-to-industry transfer shows what is possible when these conditions are met. --- ### Enterprise AI & Open Source: What Three Practitioners Said About a Sustainable Architecture URL: https://hendorf.com/en/blog/enterprise-ai-open-source/ Published: 2025-04-20 Tags: enterprise-ai, open-source, ai-strategy, digital-sovereignty, python-adoption, regulated-industries ## At a Glance In the fireside chat *"AI in Reality: Enterprise AI & Open-Source Innovation"* at PyCon DE & PyData 2025, three practitioners sat on stage who work productively with AI in very different contexts: Walid Mehanna (Chief Data & AI Officer, Merck), Dr. Alexander Beck (six years CTO and Equity Partner at Quoniam Asset Management) and Ines Montani (Co-Founder and CEO, Explosion / spaCy). The discussion was not about hype but about what works in the engine room of large and small organisations, and what does not. Four lines ran through the hour: - In enterprise contexts in 2025, open source is the pragmatic default, not the idealistic exception. It lowers vendor dependency and accelerates innovation. - Standardising on a single stack (Python) breaks down communication barriers between teams and accelerates delivery cycles, but costs friction at the start. - The German and Nordic Mittelstand often moves faster than US corporates with a start-up coat of paint, because decision paths are flatter. - Regulation isn't the actual problem; unclear regulation is. Anyone who doesn't know what *"good enough"* means builds paralysis instead of safety. ## Merck: Pivoting as a Corporate Capability Walid Mehanna, Chief Data & AI Officer at Merck (a 357-year-old family business in its 13th generation, with 63,000 employees worldwide), described his organisation with a punchline that says a lot about the culture: *"If Merck has one ability, it is pivoting."* From pharmacy to pharmaceuticals to life sciences and electronics: change is part of the operating repertoire, not something that lives in the strategy workshop. Mehanna's AI strategy rests on what he himself calls the *"holy trinity"*: *people first* (mindset and skillset), *ways of working* (orchestration and integration across teams), *technology* (libraries and services). That sounds simple, but at 63,000 people it isn't. His stated ambition: *"My dream is that you don't need me anymore because 63,000 people in this company breathe data and AI, everybody knows what it is, when it's useful, when it's not useful and can apply it."* That is not marketing, it is an ecosystem approach: security, compliance and cost efficiency combined with broad-based education, from low-code tools for domain experts to full-stack development inside the engineering teams. Notable was Mehanna's stance on coding standards: *"freedom in a frame"*. As long as work is properly documented and run-time costs don't explode, he has no interest in monolithic mandates. He expects modern, AI-assisted IDEs to converge consistency and quality on their own. Two Mehanna stories illustrated his pragmatism. First, the MyGPT episode: when ChatGPT went public, he wanted to follow suit. His team drily informed him that the NDA had been signed nine months earlier. Six months later, an in-house version was running on Microsoft Cognitive Services. Later, the team decided to retire its own product in favour of a partnership with the Berlin start-up LangDoc. Mehanna's reasoning, in his own words: *"They have four developers and they're a startup. They're working day and night. They don't have family. ... They will overtake us in three to six months. So why not join forces now?"* That is lived pivot culture, even against your own code. Second, the sandbox logic: in a global group, there is always a jurisdiction in which an experiment is regulatorily viable. *"But in essence, it's piloting in a sandbox"*, Mehanna summed up. What proves itself there then scales, with clarity about the risks, into other markets. ## Quoniam: Standardisation as Cultural Revolution Dr. Alexander Beck spent six years as CTO and Equity Partner at Quoniam Asset Management, a quantitative asset manager with more than €24 bn assets under management, part of the Union Investment Group. He described the migration of a fragmented stack: *"Back then we had SAS, R, C Sharp, a little bit of Python, a lot of SQL. And we harmonized this into Python."* In a regulated financial context (BAIT, [DORA](/en/glossary/dora/)), this was not a small decision. The effect was measurable. Where application development had previously run in C# and search in R, Beck described the prior state plainly: *"The two teams couldn't talk to each other. They couldn't learn from each other. They couldn't share code."* Today research and application development work on the same code. Domain experts understand what the engineering teams are building, instead of having to delegate it. Beck's honest assessment: *"The first steps everyone has to realign a little bit their way of working, that causes friction. ... But once you are behind these initial fires, you really feel how this harmonization speeds you up and makes you better."* Standardisation is not sexy, it works. Asked about the best decision of his time as CTO, Beck answered without hesitation: *"This really goes back to 2020, where we said for the whole company: We are a Python company."* Beck added two further operational points. First, the reality-check number: at around 100 employees, Quoniam has *"at least five very good use cases ... maybe ten more in the pipeline that have potential"*. That is an order of magnitude that is more honest than boardroom slides with three-digit use-case lists. Second, the LLM sandbox architecture: APIs to Azure OpenAI inside a contractually secured frame: the *"safe haven"* in which productive use cases can emerge without leaving the regulatory perimeter. [See related Whitepaper](https://hendorf.com/de/case-studies/transformation-asset-management/) ## Explosion / spaCy: Open Source as Business Model, and the Mittelstand Punchline Ines Montani, Co-Founder and CEO of Explosion and core developer of spaCy, brought the perspective of the open-source vendor into the panel. Her business model is counter-intuitive: if the spaCy documentation were worse, no one would use spaCy; if it were too good, there would be no consulting business. Explosion does not earn money by explaining how to use spaCy to teams, it earns money on workflows and on building proprietary models in-house. Montani described the truly hard work like this: *"How to take a business problem and break that down into components that you can solve with machine learning."* One of her most important admonitions, the kind that lands in any boardroom discussion: *"This is not a competition. You're allowed to make problems easier."* Montani's most valuable observation, and a corrective to the usual picture of US innovation versus German inertia: the most agile early adopters were not US start-ups. *"Even Mittelstand companies in the Nordics, they have, you know, sort of a slightly different mentality. We saw those first before we saw the German companies."* The reason is not regulation or hype culture; it is flat decision paths. The counter-punchline is an anecdote that lands instantly in board conversations. Montani described US corporates *"that operate more like startups and pride themselves on like how they've disrupted whatever"*, and fail on a $2,000 licence. In the original: *"They were unable to buy a license from us because ... spent weeks in legal and eventually failed for ... something that cost a few thousand euros, they were unable to purchase it because we refuse to just change the jurisdiction of our contract for a purchase of $2,000."* That is the actual vendor-lock-in effect: not in the tool but in the procurement and legal architecture. Montani was equally clear about prestige projects: *"Some companies actually want prestige projects, and they usually fail. It's like, oh, let's just do a chatbot. And that never makes it out of the prototype phase."* Successful, she said, are the teams that recognise the truly hard work isn't model training but the question of how to break a business problem into solvable components. A second, similarly tidy line: *"We've learned outsourcing development never really worked. We learned that in the 90s. Outsourcing data development, annotation. We've also realized that doesn't work."* In-house development with domain experts at the table is the recurring success formula. ## The Regulation Paradox Beck and Montani converged on a diagnosis that is central for regulated industries. Beck, from the financial-sector perspective: *"In Europe there is too much weight on regulating things."* The problem is not regulation itself but its lack of definition: *"When the regulator itself cannot really tell you what they expect, you're kind of left alone on high sea, and at some point in the future, somebody may or may not come and tell you, oh, all what you did is wrong."* The consequence: more paper, less value. Montani brought the counterpoint that often goes missing: *"We need to understand the technology in order to regulate it well."* That is the second half of the diagnosis. Anyone who is responsible for code reviews without being able to read code creates exactly the paralysis Beck describes. Mehanna added from the corporate perspective: the talent base in Germany and Europe is solid, the industrial application fields abundant. What is missing is scaling mechanics. Start-ups that want to grow are often told they have to go to the US to find growth and customers. Mehanna's reading: *"We can learn from the US, we can learn from China, but we have to find our own way."* ## Make or Buy, and What Standards Cost Today On the question of coding guidelines, all three agreed: 2025 looks different from four years ago. Where long rulebooks used to be written, the answer today is simply: *"Let's use UV, let's use Ruff."* Great open-source projects have built great standards that are easy to follow. Hours that used to flow into style debates are freed up for substance. Beck added that code reviews and vulnerability scans are of course still part of the process: standards do not replace discipline, they make it easier to enforce. For Mehanna, the [make-or-buy question](/en/boardroom-to-code-s26/) wasn't a matter of belief. *"Our strategy is find the best solution that works for the company ... and scale it in a global environment."* On selection: *"If it's open source, great. That is always our preferred go-to when there is something that is enterprise ready."* If not, look for partners: *"partners can be big or small. If they're small, even better, because we can offer them more"*. The other side, large multinationals like Merck itself, *"only want our money and our brand to a certain degree"*. ## Beyond Hype To close, a question that is rarely answered honestly: which technology, beyond the noise, should you really be watching? Mehanna placed his bet on the integration of hardware and software: *"I still believe there is a beautiful competitive advantage in merging hardware and software. ... Apple has played very nicely. More and more companies like Groq in the inference space ... Doing an integrated view is a heavy investment, but I believe it's also a high risk, high opportunity play. I would love to see a European player doing this."* This is not academic geopolitics, it is the answer to where [inference cost](/en/glossary/inference-cost-tco/) advantages will come from in five years. Montani put the product lever centre stage with an image that sticks. In the past, she said, people used to wake others up by knocking on windows. *"We didn't build window-knocking machines. We built alarm clocks."* The point for AI investments: anyone who replaces human tasks one-for-one is thinking too small. The valuable innovations come from redefining the task itself. Beck came at it from a third angle: in a world of growing geopolitical fragmentation, every European, and especially every German, offering deserves a close look. That is also a make-or-buy statement, just with a different sign. A concrete Mittelstand example from Mehanna for the thesis "AI has long since arrived in German industry": Trumpf uses AI to optimise cutting layouts on sheet-metal machines and to reduce waste. No generative drama, simply a better optimisation layer over an old machine. ## So What Three lines from the fireside chat that show up again in every advisory discussion about regulated industries. ### First: Standardisation The unsexiest high-performance investment of the next five years. Quoniam's Python migration shows that the friction at the start is real, and the speed gain afterwards is equally real. Anyone who bets on hybrid stacks because *"every team should be able to choose"* is optimising for local comfort zones rather than for delivery capacity. ### Second: Open Source In 2025 the pragmatic default in enterprise contexts, not the idealistic exception. Merck's line, when an enterprise-grade option is available, then open; otherwise partner, smaller rather than larger, is the operational translation of that stance. ### Third: The Mittelstand More agile than the usual innovation narrative claims. The real brake is not German inertia but procurement and legal architecture, and that holds in the US at least as clearly as it does in Europe. {{cta}} --- ### Data Quality Assurance: The Foundation for Reliable AI Systems URL: https://hendorf.com/en/blog/data-quality-assurance/ Published: 2025-03-15 Tags: Data Quality, AI, Data Quality Management, ML Operations, Data Governance Data quality is not a tooling problem. It is a strategic decision, one that determines how far your AI transformation can actually go. *Welcome to the Wild West!* That is a fair description of the current state of play around data quality. Everywhere people *still* preach "Data is the new oil", and yet years of work with companies have taught me one uncomfortable truth: **not every barrel of crude can be refined.** Many organisations collect data wildly without ever building a working refinery. They have volume, but no reliability. And then they wonder why their AI systems fail, not on the algorithms, but on the data underneath. This article makes the case that Data Quality Assurance is not optional but the foundation of every successful AI system. The point is strategic: how do I build data quality in from the start, rather than as repair work at the end? > **Context:** This piece summarises my talk at *Science Sparks Start-Ups* at Heidelberg University, a transfer format that mirrors enterprise consulting experience back to research-driven spin-outs. The point is not academic: what large organisations have learned the hard way over years, young spin-outs can avoid before it becomes a structural mortgage. The observations come from advisory mandates inside established companies: the audience is founders who can set up their data architecture properly now, instead of expensively repairing it later. ## The Inconvenient Truth About Data in the Real World Students learn with perfectly clean datasets like Iris or Titanic. Inside real companies it looks different: a chaos of incomplete, inconsistent and outdated data that has grown over decades. These are not isolated cases. This is the norm. In almost every company I have advised, **at least one so-called "ghost field" existed**: a data field whose meaning no one remembered, but which was still being used in critical processes. In one case it was a field that had been intended ten years ago for a specific business process. That process had long since disappeared, but the data was still being captured. The problems I most often run into: **incomplete values**, which force AI models to invent the gaps. **Inconsistent formats**, where an address sits in one field one time and three the next. **Outdated data** that no one updates because no one is responsible. **Missing standardisation**, with each department maintaining its own conventions. **Data silos**, where customer data lives in five different systems. And underneath all of it: **poor data governance**, no clear responsibilities, no processes, no documentation. ## The Anatomy of an AI System Every working AI system (whether Amazon Echo or a recommendation engine inside a bank) rests on an invisible foundation: **technology is just the tip of the iceberg.** Underneath sits a vast mountain of data sources, human work and complex dependencies. In enterprise environments this problem becomes exponentially harder. Data does not come from one source but from people (manual entries, user feedback), machines (sensors, APIs, log files) and legacy systems that have grown over decades. Each system has its own structure, its own data types, its own sources of error. The classic silo problem: data is distributed decentrally, every system has its gatekeepers, integration is error-prone, and because no one was ever responsible for the whole, an inconsistent data governance regime emerges. The result is an organisation that has volumes of data but no idea how the pieces connect. ## Data Quality: Definition and Dimensions Data quality is not abstract. It is the fitness of data for its intended use. Concretely: a customer list with 50,000 entries is useless if 30 percent have incomplete addresses or are duplicates. The essential dimensions: ### Accuracy The first question I ask: do the data correspond to reality? Not "does it look right", can I verify it? I have seen systems in which names were deliberately falsified (for customer protection), and then analyses were built on top of them and led to entirely wrong conclusions. ### Believability Can the people involved trust the data? That sounds soft, but it is business-critical. If a sales rep does not believe the customer data in the CRM is correct, they fall back on their own network, and you lose the integration. ### Completeness In AI systems this is particularly critical. Missing values do not just lead to imprecise forecasts; they cause models to systematically hallucinate around the gaps. A bank that has not captured income for 20 percent of its customers trains a credit-risk model that implicitly learns "missing information equals low risk". ### Consistency Across system boundaries, this is often the biggest problem in large organisations. A single customer can exist in five different CRM systems with five different addresses. A transaction is captured with a timestamp one day, with only the date the next. ### Timeliness Nuanced. Not all data has to be real-time fresh; some can be 30 days old. But you have to know how old your data is, and whether that is acceptable for your use case. ### Validity and Uniqueness Round out the picture: do the data conform to defined rules? Are there unintended duplicates? ## AI-Specific Data Quality Problems In AI systems, poor data quality becomes an existential hazard. This is not like a faulty Excel sheet; it is like an aircraft with a leaking fuel tank. ### Bias in training data The classic problem: a credit-decision model trained on historical data in which certain groups were systematically disadvantaged will perpetuate those biases, and possibly amplify them. This is not just ethically problematic; it is regulatorily high-risk. I have seen companies forced to pull their models from the market because bias was only detected after they went live. ### Missing representation More subtle: if you train a credit-scoring model on data that is 95 percent white male customers, marginal groups are disadvantaged from the outset. The model performs statistically beautifully, on the majority population. ### Data drift The reality in production: the world changes. Customer demand shifts, economic conditions change, your competition reacts. The model that worked perfectly in 2023 delivers wrong predictions in 2025, not because something broke, but because the data distribution has shifted. ### Scalability problems Surface when you scale from 100 transactions a day to 100,000. What works in test breaks in production, not always because of the code, but because data quality issues become exponentially visible at volume. ## The Typical Anti-Pattern: Technology Before Foundation > The shiny AI in the front end, the crumbling data infrastructure in the back end In my advisory work I keep seeing the same pattern: **shiny AI in the front end, crumbling data infrastructure in the back end.** Boards want to tick the "AI" box, and quickly commission a machine-learning project. No one asks: are our data even fit for this? The result: large investments in tools, infrastructure and talent, but the system does not work. Because the data does not allow it. Then the next typical reaction: "We need better tools!" So another data-quality platform gets bought, expensively. But without governance, without responsible owners, without documented standards. The tool sits unused while data quality continues to be a spot problem, repaired in the next emergency. ## Tooling Sprawl: The MAD Landscape Problem > Tools are just tools, not saviours. The **MAD (ML, AI & Data) Landscape 2024** is, in a word, overwhelming. Hundreds of tools for data quality, governance and monitoring. A constant stream of new VC-funded start-ups, each promising to "solve" data quality. Everyone claims to be the answer. This is exactly what I call "technology sprawl" on the home page. And it is the big trap for companies: the technology is not the problem; the missing strategy about which technology is even needed is. I have worked with organisations that ran five different data-quality tools at once, because each layer (data engineering, analytics, ML) had its "right" tool. None of them was wrong. Together they were a nightmare. On top of that: demand for genuine data-quality experts is huge. But few have real experience in production systems at meaningful scale. Even fewer understand the combination of technology and organisational data culture. In many organisations, data quality remains a lonely mission: one or two people fight windmills while the rest of the organisation assumes "quality just happens". ## Strategic Approaches: From Start-Up to Enterprise The path to reliable data quality differs by organisational size, but the principles are the same. ### Mindset **Quality by design, not repair.** This is the core difference between organisations where data quality works and ones where it keeps breaking. Quality by design means: data quality is not something that comes "later". It is part of the system from day one. When a new customer record is created, the validations are already there. When a new data source is integrated, data lineage is documented. ### Monitoring and observability Are then the early-warning system. A data catalogue gives an overview of what data exists and where it comes from. Automated quality checks run continuously in the background. Alerting notifies you when problems emerge, not when they have already had business impact. ### Governance has to be proportional Start-ups do not need the same governance as a financial institution. But there must be clear responsibilities: who owns which data? Which standards apply? Documentation (what do the data mean, how do they come about) is often neglected and is a thousand times more expensive later. ### Data Technically that means: data structures must be able to evolve (schema evolution). Incoming data must be validated. The origin of every data point must be traceable (data lineage). That sounds complex but is relatively manageable on modern platforms. ## Concrete Steps: By Level of Maturity For **start-ups**, the advantage is clear: you can do it right from the start. Define data-quality criteria for your use case. Implement validations from day one. Document data sources and their meaning. Small efforts at the start that save you years. For **established companies** it is harder, but not impossible. A data-quality assessment shows you objectively where you stand. That is not pleasant. I have run assessments that classified 40-60 percent of the data as "problematic". But without diagnosis there is no cure. Then you need stakeholder alignment: the CFO, the CIO, the head of analytics all have to understand why this matters. That is more political than technical. Then comes [incremental improvement](/en/case-studies/data-strategy-infrastructure/), not everything at once, but systematically tackling the highest risks first. And finally: tool consolidation. Fewer tools, used properly and actually adopted. **For everyone:** data quality is not optional. It is as fundamental as security or compliance, except that most companies have not understood that yet. Prevention is cheaper than cure. And the most important point: **people over technology.** The best data-quality platform is useless if no one actually maintains, documents and takes responsibility for the quality of the data. ## Outlook: The Role of Automation AI will increasingly automate data quality: anomaly detection, outlier detection, even automatic data cleansing will get smarter. But the strategic questions remain human: **Which data is critical?** Only the company can answer that. **How bad is acceptable?** A data-quality level of 95 percent is acceptable in a ticketing application; in credit scoring it is not. **Which costs justify which improvements?** That is a business decision, not a technical one. And the most important thing remains: **culture.** If it is not embedded in an organisation that data quality matters, even the best technology will not work. ## Conclusion: Foundations, Not Firefighting This is the central point: **data quality is not something you do on the side or "fix" with a tool.** It is a strategic decision that affects your entire AI transformation strategy. The advisory approach I have developed over years comes down to this: **foundations, not firefighting.** That means we do not start with "let's quickly train an AI model"; we start with the uncomfortable question: "are our data even fit for this?" If not (and in most cases they are not), we have to fix that first. That sounds expensive and time-consuming. In the short term it is. In the long term you save massively, because you do not constantly have to repair systems whose foundations are crumbling. {{cta}} --- *This article is based on my talk at Science Sparks Start-Ups at [Heidelberg University](https://www.uni-heidelberg.de/en) (a transfer format for research-driven spin-outs) and on years of advising organisations through their AI transformation. The patterns described here come from enterprise mandates; the briefing for founders is: don't repeat these mistakes, avoid them structurally from the start.* --- ### Overfitted Promises: AI in Coding Research – Hype vs. Evolution URL: https://hendorf.com/en/blog/overfitted-promises/ Published: 2025-03-15 Tags: AI Coding, Hype vs Reality, Newton, Alchemy, Development Tools *As a technology leader, who can you actually trust? What works with AI coding tools, what doesn't, and what does that mean for your engineering organisation?* ## At a Glance - **The thesis**: today's AI coding tools work more like alchemy than like established science: trial and error with black-box outputs. - **The good news**: that isn't pejorative. Newton's alchemical research led to modern chemistry. - **The strategic point**: you have to understand where AI genuinely makes you productive (boilerplate, bug fixes, documentation) and where it fails (architecture decisions, complex security, domain-specific logic). - **The consequence**: not every team and not every task is equally suited to AI coding tools. Adoption needs strategy, not just enthusiasm. --- > **For context:** This piece is an opinionated vision piece, not a playbook. The technology moves so fast that it would be dishonest to lay down definitive best practices here. What you are reading is a hypothesis, and above all an argument for developing fresh evaluation frames, instead of relabelling existing software engineering routines as "and now with AI". The tools, the language, the use cases and the boardroom expectations are new. The lens we use to judge them ought to be new too. Anyone reading this critically and thinking along is in the right place. Anyone expecting a finished roadmap will be left dissatisfied. That is by design. At neurons&neckar 2025 I put forward a provocative thesis: **today's AI coding tools resemble alchemy more than modern science**. But this is neither pure criticism nor hype. It is a sober stocktaking, and a look at how your engineering organisation should deal with it. ## Newton: Between Alchemy and Revolution Isaac Newton wrote more than a million words on alchemy, most of them unpublished in his lifetime. He was looking for the philosopher's stone, to transmute metals and to attain immortality. He experimented with "Diana's tree" and believed metals could "grow" and had life-like properties. His verdict: he missed the goal. The consequences, however, were enormous. His alchemical experiments were methodical and meticulously documented. They led to scientific breakthroughs in chemistry, optics and thermodynamics. **Newton needed alchemy in order to invent science.** That is the pattern I see today around AI coding tools: trial and error with an invisible mechanism, but not without value. ## The Parallel: AI Coding Today The problem I observe in companies: they roll out GitHub Copilot or Cursor, expect a 30 percent productivity jump, and are then puzzled when reality turns out to be more complex. Today's AI coding tools work like alchemy, not like established science: - They want the right words (prompt engineering), but the rule is invisible. - They produce astonishing results on simple tasks, and fail on subtle requirements. - Why a prompt works or doesn't cannot be explained systematically. - You have to experiment, adjust, retry: trial and error. This is not meant unkindly. It describes today's reality: **AI coding tools are not yet scientifically predictable enough to integrate blindly into critical systems.** We trust compilers too (program code → machine code), but those are strictly deterministic. > AI coding is a new layer that we still have to learn how to handle and how to fence in. But there is a pattern. And once we understand the pattern (once we know where these tools work and where they don't) we can deploy them deliberately. ## Where AI Coding Tools Really Work, and Where They Don't The good news: there is a pattern. Not every task is the same. ### Where they consistently deliver - **Boilerplate and standard patterns**: repetitive structures that look the same everywhere. - **Code completion and suggestions**: intelligent auto-completion at 80 percent similarity to training data. - **Simple bug fixes**: obvious mistakes (typos, wrong function calls, simple logic errors). - **Documentation and explanations**: turning code into prose; explaining concepts to beginners. - **Trivial conversions**: translating code between similar languages (Python ↔ JavaScript). ### Where it gets alchemical, and you need to pay attention - **Architecture decisions**: when you need microservices versus a monolith requires business understanding, not code generation. - **Security**: vulnerabilities are often context-dependent. A generated token handler can look secure without being secure. - **Performance optimisations**: subtle improvements need system understanding; AI can suggest variants but not understand the why. - **Domain-specific logic**: banking rules, compliance, business logic: none of this can be fully grasped by an LLM from the outside. - **Debugging complex problems**: when ten systems hang together and one isn't working, you need systematic thinking, not pattern matching. This is not a defect of the tools. It is their boundary. **The question for you as a technology leader is: which of these tasks dominate inside my team?** ## A Strategic Example: The Agentic AI Problem A second pattern shows up around AI agents (extended discussion in the post "[Agentic AI](/en/glossary/agentic-ai/)"). Automated systems meant to act on their own need clear guard-rails. It is the same problem: humans have to define the limits of where the AI is allowed to act and where it is not. The AI itself will not get this right without human guidance. The same applies to code generation: **an AI cannot decide for itself whether generated code is "good enough" for production.** It can iterate quickly, but the final call (the critical call) remains with humans. ## The Evolution: From Alchemy to Science Newton needed 50 years to get from alchemy to science. We will be faster. But we are not there yet. The direction is clear: - **Better explainability**: tools that show why a suggestion works (or why it doesn't). - **Specialisation**: not one general LLM for everything, but specialised models for SQL, security, test writing and so on. - **Reproducibility**: not a different result each time for the same query. - **Measurement instead of feeling**: clear metrics for productivity uplift, not just "feels faster". - **Systematic quality assurance**: not hoping the generated code is good, knowing it. - **New [evaluation](/en/glossary/evaluation/) frames**: the question "does the code work?" is not enough. We need methods that evaluate co-evolving human–AI systems. Software engineering reviews from the pre-LLM era are a stopgap lens, not the right tool. **This transition is your responsibility as a leader, not the tool vendors'.** You have to plan today when and how you adopt these tools, and where you don't. ## What This Means Strategically: Four Heuristics for the Transition These four points are not best practices: the technology is too young, the material to distil best practices from is not there. They are heuristics that hold up in today's phase and deserve to be reassessed in twelve months. ### 1. Map, don't evangelise Before you adopt AI coding tools, you have to know: which tasks dominate in our work? Are they 70 percent boilerplate (large opportunity) or 70 percent complex domain logic (limited opportunity)? No one-size-fits-all: different teams need different strategies. ### 2. Start with small pilots Begin with a team that does a lot of boilerplate and bug-fix work. Measure: what becomes faster, what doesn't? Which bugs appear? Where do you need more reviews? Then scale only if the balance is positive. ### 3. Quality gates are not optional Code from AI needs reviews just as good as any other code, only different. Not "does it work?", but "is the approach safe?" and "have we done it like this before, or is this new?". These reviews have to be done by your architects, not by junior developers. ### 4. Fence in the risk Some codebases are too critical for AI experiments. Security, payment processing, core logic: AI-generated code does not belong there until the tools are more mature. That is not technophobia; it is risk management. ## The Uncomfortable Truth, and the Opportunity Today's AI coding tools are **overfitted to their current applications**. That means: they perform brilliantly inside their training corridor (boilerplate, APIs, simple patterns) and less well the further you move away from it. This is not malicious. This is physics. **But:** Newton's alchemy was also "overfitted" to the transmutation of metals. And yet it led to modern chemistry, to thermodynamics, to optics. Not because the original question was right, but because he experimented and learned systematically. That is what is happening with AI coding today. We are in the "alchemical" phase. We are in the middle of discovery. The companies that understand this now (that encourage their teams to experiment systematically, that introduce quality gates, that learn where these tools belong and where they don't) will reap the benefits when the science arrives. The others will later wonder why they lost so much time. ## Conclusion: Tell Hype from Opportunity I often hear [in boardrooms](/en/boardroom-to-code-s26/): "We have to invest in AI" or "Everyone's already using Copilot." That is disorientation in the hype. The right question is not "do we need AI coding?" but "where does AI coding pay off economically in our organisation, and where does it add risk?" The answer is different at every company. And if you are unsure (if you do not know how to make that assessment, which tools fit you, how to roll them out without creating chaos) that is exactly what I am here for. **Not every AI promise is a promise. Some are real opportunities. The difference is strategy, and the courage to look at the new with fresh eyes, instead of forcing it into the old lens.** {{cta}} --- ### Why AI Projects Fail URL: https://hendorf.com/en/blog/why-ai-projects-fail/ Published: 2024-11-15 Tags: AI, AI Strategy, Transformation, Management, Organisation > **Extended web version of the article published in [Red Stack Magazin 03/2025](https://my.doag.org/download.5c1302a8a90f388ff792b1c551b1a002_1/) (in German), based on the talk at the KI-Navigator Conference 2024.** AI projects rarely fail on the technology. They fail on people and organisations: on conflicting expectations, missing strategy, and technology sprawl without governance. This piece walks through the patterns that most often stall programmes in regulated industries, and the levers that successful programmes pull. ## At a Glance **Why it matters:** AI projects fail systematically, but not on the technology. Knowing the actual root causes (strategy gaps, organisational sprawl, missing governance) lets you address them deliberately and protect your investment. ### Key takeaways - AI projects fail on unrealistic expectations, strategy gaps, and a leadership layer that rarely comes into contact with the technical depth. - Technology sprawl and vendor lock-in are predictable traps, not unfortunate accidents. - A stable foundation of data strategy, organisational maturity and open standards is decisive. - Open-source first is, in regulated enterprises, the strategically more sustainable choice, for auditability, [sovereignty](/en/glossary/sovereign-ai/) and total cost of ownership. ## The Three Acts of AI Failure ### Act 1: The Diagnosis Anyone who has accompanied AI projects for years sees the same picture: down in the engine room the technology works. Up on the bridge the compass is lost. The causes almost never sit in the models. They sit with the actors who arrive at the same project with different expectations and different agendas. Media and tech evangelists announce daily breakthroughs; even good journalists often lack the technical depth, because sensation rewards the click and differentiation does not. The result is a hype noise that gives boards the feeling that they urgently have to do something, without it being clear what. Leadership reacts to that pressure with the fear of falling behind, and expects quick wins that aren't realistic. There is a structural side to this in Germany: in software-native companies the engineering background of the founders still shapes today's executive layer. In long-established German corporations, the typical career paths (sales, finance, legal) rarely route IT experience all the way into the boardroom. That is not an individual reproach; it is a structural question. The consequence is uncertainty in handling technical depth, and with it, susceptibility to tools that look impressive but don't carry weight. AI consulting, finally, knows the latest technologies, and has to sell them. It sees the gap between reality and customer expectation, and at the same time stands under enormous technical pressure: the field of language models alone has been completely overhauled three times in the past twelve years: from Word2Vec and GloVe through the transformer architecture and BERT to today's large language models like GPT-4 and Llama. Anyone who has not personally rowed those waves either sells outdated architectures or the next vendor promise. ### Act 2: The Patterns In advisory practice, three patterns show up over and over, independent of industry and company size. The first is the prototype dilemma: six months of development, three years of discussion. The prototype works in the lab. No one thought about the production environment, real-world data quality looks completely different, organisational processes don't fit the technology, and change management was never part of the project from the start. What is left is a demo video and a roadmap nobody believes in any more. The second is the silo solution: marketing builds on Tool A for content generation, IT develops Tool B for data analysis, HR pilots Tool C for recruiting. Every department has its own vendor, its own data definition, its own contract. What emerges is technology sprawl without shared infrastructure: fragmented data flows, no scalability, and a compliance landscape no one can survey any more. In regulated industries this pattern is particularly expensive, because every island generates its own audit trail. The third is vendor lock-in. Many companies bet on proprietary cloud services without asking the uncomfortable questions: what happens at the next price increase? How do we get out if the model degrades or the terms change? Where does our data ultimately sit, and what is lost in the worst case? ### Act 3: The Success Factors Successful AI implementations look surprisingly similar. Three axes run through almost every transformation that has worked in my experience. On the people axis, sequencing decides. Stakeholders are involved early, from tech teams to C-level. Expectations are set realistically: what AI can do today, what it cannot. Change management is an integral part of the programme, not an afterthought activated three weeks before go-live. On the technical axis, an open-source-first architecture carries the load. It creates flexibility and control, lets new models be integrated modularly, and provides the auditability that is non-negotiable in regulated environments. The longer argument for why open models are, in enterprise contexts, the economically and regulatorily more grown-up choice is carried by the companion post *[Stop Waiting, Start Shipping](../stop-waiting-start-shipping/)*. Add to this quality assurance from day one: [evaluation](/en/glossary/evaluation/) and monitoring are set up with the first model, not after the first incident. On the organisational axis, data strategy comes before tool shopping: what do we have, what do we need, at what quality. Governance is clear: [who decides what, when and how](/en/boardroom-to-code-s26/). And the approach is iterative: small steps, fast learnings, honest corrections. ## What Distinguishes Successful AI Transformations After more than ten years in the data science community and hundreds of project contacts, I see one factor that does not come out of methodology textbooks: the most successful AI implementations emerge where knowledge flows across domain boundaries. In the Python community I have watched since 2014 how people from completely different fields work on problems together: from the European Space Agency to fintech start-ups, from climate research to industrial computer vision. At a PyData conference an engineer from a banking team picks up a pipeline idea from astronomy, because the same class of time-series problem has been solved there for ten years longer. This cross-pollination produces better solutions, because many supposedly new problems have long since been worked through in other fields. That is exactly the job of an advisor who doesn't just know the latest tool generation, but knows which patterns transfer. ## Concrete Recommendations 1. **Before you invest in tools:** analyse what your people actually need, and what the regulatory landscape allows. 2. **Develop a data strategy** before you train AI models. 3. **Bet on open standards**: they secure optionality, auditability and negotiating position. In regulated environments, open source is the strategically more grown-up choice, not the budget version. 4. **Build internal expertise:** external advisors can show the way; you have to walk it yourselves. 5. **Start small:** a working pilot is worth more than ten failed visions. ## Conclusion AI projects do not fail on the technology. They fail on strategy gaps, organisational sprawl and missing governance. Anyone who builds a stable foundation of clear prioritisation, data-centric planning and open-source-first architecture creates the precondition for sustainable success. The technology is there. The question is: are you organisationally and strategically ready for it? --- ### The Economics of Prediction Machines: Understanding AI as an Economic Factor URL: https://hendorf.com/en/blog/economics-of-prediction-machines/ Published: 2020-11-15 Tags: Prediction Machines, AI Economics, ABB, Industry 4.0, Economic Impact ## At a Glance AI is not IT. It is research and development. That is the central insight my advisory work with German industrial companies keeps confirming. Companies that grasp how falling prediction costs change fundamental business models, and that build their prediction capacity systematically, will have a decisive competitive edge tomorrow. This article shows how the "prediction machines" economy actually works, and why strategy matters more than code. --- > AI is not IT, it is R&D! This insight from my talk at the Netzwerkforum Smartproduction at the ABB Ability™ Customer Experience Center in 2020 has been confirmed in the years since by hundreds of conversations with boards and technology leaders. In practice I see the same pattern over and over: companies that treat AI as a technology project fail. Companies that grasp that this is about predictions (and therefore about economic disruption) create value systematically. ## The Core: What Prediction Machines Really Are The term "prediction machines" is more precise than the usual "AI" rhetoric. When I talk to industrial companies, I notice quickly that they understand all kinds of things by "AI": from robotics to chat systems. The reality is simpler and at the same time more powerful: **AI systems are machines that make predictions. That is not marketing spin, that is the economics.** When falling prediction costs (driven by better hardware, cheaper data storage, open-source software, open research) become the new normal, completely new business opportunities emerge. That isn't science fiction. That is what I see happening in client projects every day. And here is the **first strategic gap**: many companies don't think in predictions. They think in technologies. They think in IT projects. They don't think about how cheap predictions change their business model. ## The Goldman Sachs Case: A Teaching Example The much-cited Goldman Sachs example still captures pithily what happens when prediction costs fall. In 1999 the bank still needed about 600 traders to handle the US equity market. Today there are two. The rest is algorithms: prediction machines that continuously forecast market patterns and react automatically. That is not just "automation". It is the reinvention of an entire business model. And I see exactly the same pattern in my client projects across Germany and Europe. In specialty chemicals, in mechanical engineering, in automotive supply. The companies that grasp fastest that their critical business processes depend on better predictions (rather than on better sensors or better machines) are the ones that win. Where do I see this blindness most often? In two places: companies that haven't yet understood where their prediction costs could fall (which is most of them). And companies that think better predictions are an "IT procurement", instead of a strategic, iterative R&D task. ## Where the Real Opportunities Sit: What I See in Industrial Companies When I work with German and European industrial companies, I often hear: "We want to put AI into our machines." That is too narrow. The real potential is not in smart sensors but in predictions about what happens around them. ### First category Supply chain and demand planning. A mid-sized mechanical engineering firm I worked with had a classic problem: too much raw material in stock in good times, too little in bad. With better demand forecasts (not deep learning, but classical statistical methods) they were able to cut inventory by 22 percent and at the same time improve supply reliability. That is prediction machines in practice. ### Second category Unstructured information. Many companies have thousands of contracts, inspection reports and supplier correspondence. Modern natural language processing makes it possible to work with these systematically, not in order to automate completely, but in order to assess risks, find anomalies and detect compliance problems early. ### Third category Finance and risk. Better forecasts of cash flow, of supplier default risk, of fraud patterns in accounting. What I keep seeing here: companies massively underestimate what is possible with simple, classical statistical methods. They want to start straight away with neural networks because they have heard that "deep learning" is fashionable. That is disorientation in the hype. The iterative path is: simple statistical methods first, then Bayesian methods, then (only if needed) the large models. ## Why Most Companies Still Fail The strategy gap is real. Companies see that "AI" matters. They start projects. But then they fail at the same point, because they have not understood that **AI is not IT, it is R&D**. That is a fundamental difference. IT projects have clear requirements. You write a spec, commission a vendor, and they deliver a system that works or doesn't. R&D is different. In R&D, uncertainty is normal. Experiments fail. You have to iterate. That requires different governance, different budget models, different communication with the board. I see again and again companies that budget three million euros for an "AI project", set up a large task force, and after six months realise the biggest blockers aren't technical but organisational. Data access is hard, because the data is spread across various legacy systems. The line of business doesn't understand why a model only has 75 percent accuracy. IT governance won't allow open-source software in production. These are not technology problems. They are strategy problems. And they have to be solved by the board. ## Best Practice: How It Works From my experience with successfully implemented prediction machines in German companies, there is a proven pattern: ### Phase 1: Definition (weeks 1-2) Don't start with code. Start with the business question. Which predictions are truly critical for our business? Where do we already save money today through manual "predictions" (i.e. well-informed guesses)? Where do we lose money because our predictions are wrong? ### Phase 2: Exploration (weeks 3-8) Experiment with the data you have. Simple statistical methods. Quick prototypes. Two or three experts. Not 20 people in a task force. The insight "the model could deliver 18 percent better results, but we need more data" is more successful than a large project that fails because the requirements were never clarified. ### Phase 3: Iteration (weeks 9-24) The working model gets gradually integrated into real processes. This is the critical phase. Don't start at 100 percent accuracy. Start at "just better than the status quo". Real people use the system, give feedback. The model is continuously calibrated. > The take-off is not the goal. A safe landing is the goal. > Behind this sits a core principle: **the take-off is not the goal. A safe landing is the goal.** Many projects look beautiful at the demo stage and then fail miserably at production rollout. Success is not measured in accuracy scores but in: does it save real people time? Are business decisions actually better because of the better predictions? ## The Strategic Opportunities When prediction costs fall, opportunities open up that didn't exist before. Predictive maintenance (from reactive to proactive) is one example. Better inventory optimisation is another. Mass customisation (scaling personalised solutions) becomes economically viable when predictions about customer preferences become cheap. But (and this is the most important observation from my projects) these opportunities don't appear automatically just because the technology gets cheaper. They appear when a company also anchors this new economics strategically. In other words: when the board doesn't only fund the technology but [creates the organisational conditions](/en/interim-ai-leadership/) in which R&D actually works. That is foundations, not flailing. That is the differentiator. A company that starts with real strategy ("we want to be the prediction leader in our market, and that is a 3-5-year programme") will, in five years, beat competitors who today are starting AI projects faster. ## The Central Question for Every Decision-Maker Which predictions are critical for my business model? And: am I faster at making better predictions than my competition? That is the question. Not: "do we have a ChatGPT chatbot?" or "can we put sensors on our machines?" The best prediction systems emerge when domain expertise and data-driven thinking come together. The chief data officer and the chief operating officer have to work together, not next to each other. The board has to understand that better predictions mean business models have to be rethought. ## Summary: The Economics Have Shifted The "prediction machines" economy is real. The falling cost of prediction is genuinely changing business models. This is not marketing. I see it in my projects every day. But (and this is the point) not every company turns this opportunity into success at the same rate. The strategy gap is huge. Many boards don't understand that AI is not IT. Many technology leaders think it is about better algorithms, when it is actually about better data and faster iteration. Many companies are driven by hype rather than by real strategic thought. That is disorientation in the classical sense. And it is also the opportunity for companies that stay sovereign. The best predictions don't come from the most expensive models, but from: 1. Clear business questions instead of a tools-first mentality 2. Empirical, iterative processes instead of large waterfalls 3. Cross-functional teams that combine domain expertise with data competence 4. Real R&D budgets, not IT budgets 5. Strategy from the board, not from the technology team The fact that this is so rare is exactly what will separate the strong companies from the weak ones tomorrow. {{cta}} --- ### AI for Decision-Makers: Why You Can't Buy AI, You Have to Build It URL: https://hendorf.com/en/blog/ai-for-decision-makers/ Published: 2020-03-15 Tags: ai-for-managers, data-literacy, enterprise-ai, management, misconceptions, team-culture "I need to buy AI." That sentence from a senior executive at a railway station captures the biggest misconception in management: artificial intelligence is not a product you simply buy off the shelf. After years of advising executives and data scientists, the fundamental challenges are clear, and they have less to do with technology than most people think. > **A look back from 2026:** This piece was written in 2020 and is left deliberately unchanged. The core theses (AI is not a product to buy, data quality decides, the human factor is the actual challenge) have carried through six years and one LLM surge. Some numbers and examples are visibly pre-LLM-era; I am not updating them, because the value is in the logic, not in the comparison figures. For anyone asking what long-term experience in this topic looks like: this is a trace back to the time when AI in the German Mittelstand still had to be spelled out letter by letter. ## At a Glance - AI is not a product to buy but an organisational and technological capability that you build. In 2020 this thesis still needed explaining, six years later it is consensus, but the execution is far from finished. - Success depends on company culture, clear business understanding and interdisciplinary teams, not primarily on technology. - Data literacy for decision-makers means: being able to evaluate data and make data-driven decisions, without programming yourself. - Simple solutions often beat complex AI systems. Asking the right questions remains a leadership task, in the LLM era more than ever. ## The AI Paradox in Management When you Google "AI for managers", you always see the same white robot. Google "replacing managers" and you find articles about how AI will replace executives. But Google "replacing data scientists": there the discussion is still open. For data engineers the discussion does not even exist. Which shows: the biggest threat from AI is not the technology itself but the misunderstanding around it. ### The reality of successful managers Harvard Business Review studies show: managers spend 70% of their time on administration and problem-solving. Only 10% goes to strategy and innovation, 7% to people development. At the same time they believe digital technologies and data analysis are among the most important future competencies, but massively underestimate the importance of people skills. **The decisive point**: the biggest challenge in AI projects is human communication, not technology. ## AI Isn't New, Neither Are the Problems Artificial intelligence is older than relational databases. AI was invented in the 1940s, relational databases only in the 1970s. MIT Technology Review, the magazine of one of the world's leading technology universities, was already debating the same topics in the 1980s and 1990s: "Will artificial intelligence ever fulfill its promise?" (1986, after 25 years!), "Automation" (1985), "How to keep mature industries innovative" (1987) and "Can computers create literature?" (1998). **The insight**: what you experience today as revolutionary is often a wave of developments that have been running for decades. Understanding that helps you separate signal from noise. ### Why AI works now Three factors pulled AI out of the "winter": **compute** (Moore's Law, GPU computing, specialised AI chips), **data** (the internet age, cheap storage, global data collection) and **software** (open source, global knowledge sharing, the Python ecosystem). This combination simply did not exist two decades ago. ## Data Literacy for Decision-Makers: Understanding from the Boardroom Down to the Code As a leader you do not need to be able to code, but you do need **data literacy**: the ability to evaluate data and make decisions on it. That is the core of your job as a decision-maker. Working with data has four levels. On the first two, data collection and data management, all you need to know is that professional systems handle these tasks. **Here is where it becomes relevant for you**: at the data [evaluation](/en/glossary/evaluation/) level you understand what the data means, how findings are presented, and above all whether data-driven decisions are sensible. That is your core competence. The fourth level, data application, is your main domain. Decisive here: data ethics (non-negotiable for leaders), critical thinking (which you delegate to experts) and decision evaluation (your core competence). My rule of thumb after years of advisory work: of everything your data scientists do, you may need to understand maybe 20 percent technically. The other 80 percent is change management, business understanding and people leadership. ## The 5 Biggest AI Misconceptions in Management ### 1. "Bigger is better", the Hadoop mistake Classic example: companies buy Hadoop clusters before they hire their first data scientist. Result: "we don't actually need the Hadoop cluster, we only have a few gigabytes of data." **The rule**: don't buy resources pre-emptively, understand the need first. ### 2. "Data lakes are clean lakes" Data lakes are not clear mountain lakes that you scoop clean data out of. They are complex systems made of many components, more concept than technology. Data quality comes from company culture and governance, not from technology. ### 3. "It's an IT project" Wrong. Data science and AI are **research and development**. They need an experimental culture instead of rigid processes, an open budget for experiments that can fail, interdisciplinary teams, and a willingness along the way to solve different problems than the ones planned. ### 4. "Data is objective and unbiased" Data is **always** biased. Example: a US police AI system systematically discriminates against certain groups because it was trained on historical, biased data. **The solution**: diverse teams, ethical guidelines, transparent processes. ### 5. "Deep learning solves everything" Often you can solve problems better with classical statistical methods. They are more stable and provable, explainable (often not possible with deep learning), less data-hungry, and faster to implement. ## The Garry Kasparov Principle In 1997 IBM's Deep Blue beat the world chess champion Kasparov. Many theorists had predicted: "when a machine beats humans at chess, they have overtaken us." **What actually happened?** Kasparov did not become unemployed. He says today: "computers are great tools." He developed new concepts for how humans and computers can work together. Chess is more popular today than ever. **Kasparov's quote from Pablo Picasso**: "computers are useless. They can only give us answers." **The lesson**: AI does not solve problems, it answers questions. Asking the right questions remains a human task. ## Practical Implementation: How AI Really Works ### The Titanic example The famous Titanic dataset is boring for data scientists, but for executives it is perfectly instructive: - **Simple rule**: "all women survive, all men die" = 80% accuracy - **Complex AI**: marginal improvement, but not explainable - **Domain expertise decides**: anyone who knows the story of the Titanic understands the data **Insight**: simple solutions are often better than complex AI. ### Style transfer as a teaching tool Take a Van Gogh painting, learn the style, apply it to a photo, and there is your artwork in Van Gogh style. **But careful**: it doesn't work on backlit photos. Why? Probably no backlit images in the training data. **The lesson**: AI only works as well as the data it was trained on. ### Speech synthesis, a realistic example With a MacBook and 9 days of training you can build a system that reads any English text aloud. Cost: under €1,000. **But**: 95% of research results never make it into production. Be realistic about expectations. ## Company Culture as a Success Factor ### The cooling-house experiment The psychologist Dietrich Dörner had people steer a complex system (temperature regulation). Result: under stress, people fall into erratic behaviour and maximum reactions. In a company that means: without the right culture, even the best AI projects fail. Under pressure, people fall back into old hierarchies and silos, and that is exactly what happens when an AI programme is overtaken by expectation pressure. ### Six recurring postures in AI teams In every AI team I encounter recurring patterns that lend themselves nicely to personification. Six archetypes from advisory practice, none of them wrong per se, none of them sufficient on their own: - **Anodyne Andy**: does not know the problem. He doesn't know what his team is actually working on. - **Easy Ed**: wants to look at options. Searches for alternatives before committing. - **Show-off Sarah**: wants to "do AI". The technology is the goal, not the means. - **Helpful Hannah**: has structured data and is willing to contribute it. - **Labeling Larry**: has "labelled" data. The quotation marks are important, because the quality of the labels later decides everything. - **Prepared Pam**: understands the problem and knows the NLP tool that fits it. The leader's job is not to judge these postures but to channel the energies. Show-off Sarah's enthusiasm is an asset when it lands on a problem that genuinely exists. Anodyne Andy needs clarity on the business problem before tools become a topic at all. Prepared Pam is the rare gold standard, and the profile [that carries a programme](/en/interim-ai-leadership/), if you can hire one. ### The communication principle In the Python community, astronomers can work productively with web developers, held together by an open, respectful communication culture. That experience is one of the reasons why open source, for me, is not one option among many but a structural advantage. Translated to your company: create spaces where everyone involved can honestly say "I didn't understand that", without fearing a loss of status. That is the precondition for interdisciplinary teams to work. ## What Decision-Makers Have to Do Differently **Think across industries**: innovation happens at unexpected intersections. Which problems do other industries solve in similar ways to yours? Where do you find unconventional partners? What can you learn from completely different domains? **Rethink team composition**: you don't need "AI superheroes". You need diverse teams with data engineers (making data accessible), data scientists (models and insights), domain experts (context and business understanding) and change managers (taking people with you). **Establish an experimentation culture**: Google brings only 5% of its AI models into production. Prepare yourself for that success rate. That means: budget for experiments that can fail, a learning culture instead of guaranteed outcomes, iterative development instead of big-bang projects. **Realistic timelines**: AI is a marathon, not a sprint. From experience you need 3–6 months for a proof of concept, 1–2 years for production maturity, and 2–5 years for scaling. ROI does not show up in quarters but in years. ## AI in 10 Years: The Realistic Vision Forget science fiction. AI will become part of everyday life, without our perceiving it as "AI": ### Practical applications - **Call centres**: AI analyses emails up front and gives agents context. Result: human interaction gets better, not replaced. - **Predictive maintenance**: machines flag themselves before they break. - **Document analysis**: decades of research work get automatically categorised and made searchable. ### Business models When you have an AI model that works, you can scale it horizontally. That is Google's business model: build once, use millions of times. **The decision**: do you only want to buy AI services, or do you want to develop your own scalable AI systems? ## Concrete Recommendations **Now**: develop a data strategy (where is your data? Who has access? What quality?). Map domain expertise: your best AI applications emerge where you have the deepest business understanding. Create regular formats in which technical and business stakeholders come together. **Medium term** (6–12 months): define an experimentation budget (5–10% of IT spend for AI experiments without guaranteed outcomes). Start with a pilot project: small scope, clear business value, measurable results. Implement AI literacy programmes for executives: basic understanding, not coding skills. **Long term** (1–3 years): make AI strategy part of your corporate strategy, not an isolated IT topic. Build partnerships with universities, open-source communities and unconventional industries. Establish ethics and governance before you scale. ## Conclusion: The Human Factor Decides The most important insight after years of AI advisory work: **technology is only half the rent**. Success depends on whether you solve the human challenges: - **Communication** between tech and business - **Trust** in interdisciplinary teams - **Patience** for iterative development - **Courage** to experiment without a guaranteed outcome You cannot buy AI, you have to build AI. And that is a deeply human task. **The decisive difference**: companies that understand this use AI as a strategic advantage. The others remain customers of Google, Microsoft and Amazon. {{cta}} --- ### What Two Years of Deep-Learning Experiments Teach About AI in Practice URL: https://hendorf.com/en/blog/deep-learning/ Published: 2020-02-11 Tags: deep-learning, pytorch, hands-on, strategy, enterprise-ai, practical-ai > **A look back from 2026:** This piece consolidates my hands-on experiments from 2018 and 2019, documented in talks at PyCon and PyData. I am leaving it deliberately unchanged. What you read here is the honest experience of a practitioner before the LLM wave: what worked, what failed, and why the 90-percent trap is not theory. Seven years and several model generations later, the tooling has changed fundamentally. The lessons about data quality, about the gap between research and production, and about speaking honestly about failure have not become more obsolete, only more expensive to ignore. ## At a Glance - **The 90% trap is real**: quick prototypes are easy. Production-ready systems eat 90% of the time, in my experience. - **Data quality beats algorithms**: a perfect model on bad data is worthless. - **Hype and reality live in different worlds**: online success stories are carefully picked excerpts. Backlit photos fail, German speech synthesis sounds robotic. - **Neural networks are highly specialised savants**: excellent at narrow tasks, but not intelligent in any human sense. - **Hands-on beats theory**: the apparently easy bet (generating a children's detective novel) failed. The apparently utopian bet (a narrator voice on consumer hardware) worked. Hype and reality rarely line up. - **Ask the right questions**: Garry Kasparov did not say "AI replaces us"; he said "AI is a great tool for learning." Pablo Picasso: "Computers are useless, they can only give us answers." --- ## Hands-on Instead of Hype: What I Tested Myself in 2018–2019 Between 2018 and 2019 I spent two years on intensive deep-learning experiments: from style transfer through text generation to speech synthesis. The motivation was never marketing or buzzwords. It was the concrete question: "what actually works, and what doesn't?" That experience is why I work with C-level teams today. I don't just know the theory; I know the everyday gap between sample code and production. ### The 90% trap: don't trust the first win The first success was striking. A 1980s French comic style transferred cleanly onto modern comics, on the very first try. "Oh my god," I thought, "deep learning can solve anything!" That moment repeats itself in almost every AI project: the first demo dazzles. Then reality arrives. 90% accuracy on test data becomes 60% in production. The last 10% eat 90% of the time. That is not a deep-learning problem. That is your organisation's problem. ### The data problem: models are only as good as their training data With style transfer I ran into a classic pattern: backlit photos were a disaster. The reason was simple: the COCO dataset the model had been trained on contained almost no backlit shots. The network had learned what it had been shown and could not extrapolate. That is not a theoretical problem. Every company with biased data hits the same wall: a model is only as good as the perspective of the data it was trained on. An American dataset produces American assumptions. A sales-team dataset produces sales-team assumptions. In compliance, fairness and risk discussions, this is the central question, and it is almost never solvable technically; it is an organisational one. Which discipline this implies for companies (from data lineage to governance) is laid out in the post *[Data Quality Assurance: The Foundation for Reliable AI Systems](../data-quality-assurance/)*. ### The over-ambitious attempt, and its limits I tried to generate new episodes in the style of 200 "Drei ???" audio dramas, a German children's detective series. The project fell apart quickly. Text generation produced German nonsense words like "Schloko-ljana" (an invented word, half chocolate, half proper name), but no actual stories. The chatbots were noise. That is not an anecdotal failure. It was the reality of RNNs on sequences that have to carry meaning. When I later presented these results at international conferences, the laughs from the audience came easily, and the questions at the booth afterwards were not about the technology. They were about exactly this point: when is a demo a toy, and when is it a product? It is the same question I now hear in board meetings, only without the laughs. ### Speech synthesis: the realistic potential Speech synthesis was different. With 24 hours of English audio recordings and a single GTX 1080, after nine days of training I had voices that could speak "once upon a time there was a little mermaid." Not Alexa-ready, but impressive given the resources. German speech synthesis? That revealed a different problem: there were no usable German datasets. So I built one myself. The result was understandable but robotic, and occasionally produced pronunciation slips that would have been inappropriate in a children's audio book. That is not incompetence. That is data quality. What you should take away from this: **most research never reaches production.** From my own practice, the share is somewhere around 95%, not because researchers are lazy, but because the path from "works in the lab" to "runs 24/7 under load" is a discipline of its own. ### Hands-on beats theory The overarching lesson from these experiments is mundane but decisive: **hands-on beats theory.** Generating a children's detective novel (riding the hype of the day) would have been the easy bet. Training a German narrator voice on consumer-grade hardware looked utopian. Reality showed exactly the opposite. If you want to understand the gap between marketing slide and engine room, you have to put your own hands on it. No white paper, no vendor demo and no blog post will replace that. ## Research vs. Production: An Insurmountable Gap? What I knew after those two years: research code is not production code. I saw implementations from Facebook still in Python 2.7. Jupyter notebooks stuffed with closures, where nobody but the author knew which library was responsible for what. After 38 hours of debugging I asked exactly that question. That is not a criticism of researchers. They build prototypes; that is the job. But it explains why so much research never makes it into production, and why the supposed "last step" to productionising is in fact the climb itself. ## Failure as Part of the Strategy None of this means your organisation should now reflexively spin up two dozen AI prototypes. What matters is how you steer the work: which results do we expect, how do we measure them, and how do we create safety for everyone involved? Within that frame, failure is explicitly allowed, provided it is made visible, documented and reviewed. What I see far too often in practice: flagship AI projects get pushed to "success" because only success is rewarded. A sponsor, a project lead or a team that wants to be on stage cannot afford visible setbacks, so setbacks get reframed, polished, or quietly dropped. The result is slides full of success metrics and an organisation that learns nothing from the actual project. A productive culture rewards the opposite: the people who speak openly about mistakes, dead ends and necessary course corrections. Those reports are what other teams learn from, and that is where the real value is created. In my experience, the teams that treat failure as part of their methodology are the ones that deliver durable results in the end. Not the flashy demos on the quarterly slides. In regulated industries this cultural question matters twice over: compliance and risk management depend on risks being named early. An organisation that quietly hides its AI failures internally is building exactly the weakness it is supposed to prevent in every audit report. ## What I Want to Show You These experiments were not designed to launch a start-up. They were designed to understand reality. And that is exactly what companies need now: not the marketing version of AI, but [the production version](/en/case-studies/nlp-knowledge-extraction/). If your organisation wants to introduce AI strategically (from board strategy through technical execution down to code quality) then we speak the same language. I don't just know what works. I also know where it breaks. --- ### Agile Analytics Rock the Enterprise: Open Source as a Game-Changer URL: https://hendorf.com/en/blog/agile-analytics-enterprise/ Published: 2018-03-15 Tags: Agile Analytics, Open Source, Python, Insurance, Change Management, Technological Sovereignty My client was trapped in proprietary island solutions, the worst-case scenario of typical enterprise analytics. The requirements grew, flexibility shrank, costs went up. Out of that experience came a case study that shows: **open source is not just cheaper, it is strategically superior**. And: the technological shift needs an organisational shift alongside it, in order to work. *Presented at the ZEW (Centre for European Economic Research), one of the most renowned economic research institutes in Europe.* > **A look back from 2026:** This case study was written in 2018. Back then, open source for enterprise analytics still required justification; today the question is no longer "whether" but "how consistently". I am leaving the post deliberately as it stands, because it makes three points that have, eight years on, become sharper rather than softer: vendor lock-in is an economic trap, technological sovereignty through open source is a strategic decision and not a cost-cutting programme, and change management is what tips it, not the tech stack. Anyone asking me today about AI sovereignty in regulated industries will find here the older layer of the argument that today's advisory work is built on. ## At a Glance | Aspect | Description | | --- | --- | | **Starting point** | Proprietary silos, multiple programming languages, redundancies, no big-data scalability | | **Solution** | Python-based open-source ecosystem as a shared language, with Jupyter, Pandas, PySpark | | **Result** | 50–70% reduction in development time, faster delivery, better collaboration | | **Success factor** | Parallel change management, technology alone is not enough | | **Core message** | Technological [sovereignty](/en/glossary/sovereign-ai/) through open source, not cost saving | ## The Problem: Trapped in Silos My client used proprietary industry solutions that were no longer up to the demands of the time. The starting point is typical for large companies: analysts mastered different tools (Excel, some R, C or Haskell), exchange between them was awkward, and the same processes were repeatedly reimplemented in different languages. The team wanted to deliver more analyses, but management requirements were systematically slowed by the fragmented technology landscape. The result was significant redundancy and an efficiency bottleneck that felt hard to break. ## The Solution: Python as Shared Language We rebuilt the technical architecture from scratch, not on a different proprietary solution but on **Python and the open-source ecosystem**. The reason is simple: Python has become the lingua franca of data analysis. The learning curve is shallow for analysts used to Excel, the ecosystem is mature (Jupyter for notebooks, Pandas for data manipulation, PySpark for big data, Matplotlib for visualisation) and, above all: **it eliminates the language silos and creates a shared base**. This was not primarily a cost-saving decision. It was a **decision for technological sovereignty**: independence from vendors, access to the latest developments of the global community, and the ability to adapt the tech stack to specific requirements rather than submit to a vendor's terms. ## The Implementation: Architecture and Workflow With Python as a shared language, we established an **end-to-end workflow**: data can be imported from arbitrary sources, analysis and modelling happen in the same environment, visualisation and reporting follow without media breaks. That gets rid of the export-import cycles and error sources that had previously caused expensive delays. The technical flexibility is an enormous advantage: all common data formats are supported, standard interfaces (SQL, REST APIs) work out of the box, and the solution runs where it scales: locally for prototyping, private cloud for sensitive data, public cloud for cost optimisation. **Vendor lock-in is impossible**, because there are no proprietary formats or APIs. Which also means: an exit strategy is available at any time. ## The Decisive Factor: Change Management I have to be very direct here: **technology is only half the rent.** I have seen enough failed tech projects in which the best architecture broke against a lack of acceptance. At my client, the organisational shift mattered just as much as the technical implementation. Concretely, that meant: gradual rollout (not everything at once), hands-on workshops with real examples from the analysts' day-to-day, and peer learning, using experienced colleagues as multipliers. External coaching support during the transition was indispensable. On the management side, the decisive thing was **communicating the vision, not as a cost-saving programme but as the path to more flexibility and speed**. We deliberately made early successes visible and celebrated them. We took resistance seriously (it was there, real concerns, not knee-jerk pushback) instead of ignoring it. And the parallel operation was central: the old and new solutions ran side by side for a period, which minimised the risk. Governance rules (code reviews, shared best practices, documentation) were established early on, not as bureaucracy but as the way to secure high-quality, traceable analyses. ## The Results: Hard Numbers and Soft Factors The **quantitative improvements** were impressive: development time for recurring analyses fell by 50–70%, time-to-market for new analytics products shrank dramatically, maintenance costs dropped sharply through the uniformity of the technology, and big-data requirements could be met without major new infrastructure investments. Just as important were the **qualitative gains**: knowledge exchange in the team became visibly better (a shared language creates a shared culture). Analysts suddenly had far more time for real analysis work instead of for tool wrestling. Responsiveness to new business requirements rose noticeably. And (not to be underestimated) a modern technology stack made the jobs more attractive, which made talent recruitment easier. ## Lessons Learned: What I Took From It ### First: training time is shorter than expected Excel analysts pick up Python faster than you might think. The biggest hurdle is not the technology. It is the willingness to change. When management stands behind the shift and the vision is clearly communicated, the technical retraining turns out to be surprisingly easy. ### Second: open source is enterprise-ready In 2018 that was still a real prejudice. But Pandas, Jupyter and friends are deployed by millions of users worldwide in productive, critical environments, not just in start-ups. The stability is high because the user base is so large, and error rates stay low through intensive review. ### Third: community beats vendor support Stack Overflow, GitHub Issues, specialist forums: the Python data-science community answers faster and better than the support departments of most proprietary vendors. That is a real differentiator of open source and contributes substantially to long-term productivity. ## What I Recommend to Companies If you want to take on [this transformation](/en/case-studies/transformation-asset-management/), three things have to run in parallel: a **pilot project** with manageable scope as a proof of concept, an honest **skills assessment** of the existing competencies, and a clear **governance strategy** for code, security and deployment from day one. Alongside that, you need a clear **change-management programme**: identify enthusiasts as change agents who can lead the others. Plan realistic training and budget for it, because it is an investment in your people, not a cost line. Parallel operation of old and new is not inefficiency; it is risk management that pays off. Most important: **think long-term.** Open source is not a cost-saving strategy. It is a strategic decision for independence and future-proofing. Management support is not optional; it is the precondition for teams to shape the change rather than merely endure it. ## Conclusion: From Early Adoption to Best Practice It is striking that in 2018 (the time of this presentation) we still had to justify that open source was suitable for enterprise analytics. Today this is the standard. Which means: companies that took the step back then had a competitive advantage, and they still have it. The **core advantages remain unchanged**: flexibility without vendor dependency, direct access to the innovations of the global community, attractiveness for qualified talent, future-proofing without worry about discontinued products or surprise price increases. But the **critical success factor isn't the technology; it is the organisational maturity to lead and sustain the change**. That is what counts in long-term transformations. And it is exactly where my advisory work focuses: not on the choice of technology (which is comparatively easy), but on the **strategic accompaniment of the change, the change-management competence, and the safeguarding of long-term execution**. {{cta}}