Tag: AI Automation

  • OpenAI Published a Map of AI’s Impact on EU Jobs and the Workflow Change Bucket Is Where I Live

    OpenAI Published a Map of AI’s Impact on EU Jobs and the Workflow Change Bucket Is Where I Live

    OpenAI EU AI workforce report covering job automation and workflow change in Europe

    OpenAI published Mapping Europe’s AI Workforce Opportunity this week, a report that tries to sort EU occupations by how AI will hit them. The OpenAI EU AI workforce report breaks roles into three buckets: jobs facing automation, jobs likely to grow, and the massive middle where workflows get reshaped without the role itself disappearing. That middle bucket is where I have spent the last several years of my career.

    The doom headlines will focus on bucket one. I want to talk about bucket two, because that is where the actual work is.

    What the OpenAI EU AI workforce report actually does

    The report maps occupations across the EU labour market against AI exposure, using task-level analysis rather than blunt job-title categorisation. It pulls from O*NET-style task decomposition and overlays current model capabilities to score how much of each role can be done by AI, augmented by AI, or left largely untouched.

    Three buckets come out the other end.

    The automation bucket holds roles where a large share of tasks are model-doable today. Think structured data entry, basic translation, first-line content moderation. The growth bucket holds roles that get more valuable because AI exists, including AI-adjacent engineering, training data work, and oversight roles. The workflow change bucket is the biggest of the three by headcount, and it covers knowledge workers whose individual tasks shift but whose overall job sticks around.

    The report is careful. It does not predict timelines. It does not claim to know how regulation, adoption rates, or organisational inertia will shape the actual outcome. It is a map of exposure, not a prophecy.

    Why it matters

    The workflow change bucket is the entire job description of a Power Platform developer, an RPA engineer, an automation consultant, or anyone who builds Copilot Studio agents for a living. We are the people who go into a role, decompose the tasks, and figure out which ones get handed to a flow, an agent, or a model call, and which ones the human still owns.

    The report is essentially describing the next five years of demand for this work.

    Where I think it is right: the middle bucket is huge, and most people underestimate it. Headlines about full job replacement get clicks, but the operational reality is task-level reshaping inside roles that keep their name on the org chart. A finance analyst is still a finance analyst, but half their reconciliation work now runs through an agent and they spend more time on exception handling and commentary.

    Where I think it understates the reality: the report treats workflow change as if it happens because the technology exists. It does not. Workflow change happens when someone redraws decision rights, and most organisations avoid that conversation because it is uncomfortable. I have watched plenty of automation projects stall not because the tech failed but because nobody was willing to say who owns the decision after the agent makes its recommendation.

    The report also does not capture the latency problem. A reshaped workflow where the AI step takes two seconds and the human approval step takes two days is not actually reshaped. It just has a faster front end and a longer queue.

    What I would do with it this week

    If you build automations for a living, read the report and find the occupations in your organisation that sit in the workflow change bucket. Not the automation bucket. The middle one. Those are the roles where you have the most leverage in the next twelve months, because the people in them are not afraid of being replaced. They want the boring parts gone.

    Then pick one task inside one of those roles. Just one. Decompose it. Figure out which steps a Power Automate cloud flow handles, which steps need a Copilot Studio agent, and where the human stays in the loop. Build a small version. Ship it to one team.

    This is the work. It is unglamorous. It is also exactly what the OpenAI EU AI workforce report says the EU economy is going to need a lot of, for a long time.

    I have been doing this work for a while and have written about the patterns that show up across these projects. The report does not change my day-to-day. It validates it.

    The next five years are going to be a lot of careful task decomposition inside roles that keep their names. That is fine by me.

    This post was inspired by Mapping Europe’s AI Workforce Opportunity via OpenAI.

  • Microsoft’s Intelligent Apps Post Is About Leadership and I Think They Buried the Lede

    Microsoft’s Intelligent Apps Post Is About Leadership and I Think They Buried the Lede

    Intelligent apps and human leadership in enterprise automation

    I opened Microsoft’s April post on intelligent apps and human leadership expecting another speed pitch. Faster tasks. Faster decisions. Faster output. The usual rhythm. What I actually read was a piece where the words intelligent apps and human leadership are deliberately bolted together, and I think most people skimmed past the second half.

    The leadership piece is doing all the work in that post. The intelligent app is the easy half. I keep seeing automation teams treat AI features as a productivity multiplier when the actual constraint is whether anyone in the org is willing to redesign how decisions get made. Skip that, and you ship faster versions of the same broken workflows.

    I read the post expecting another speed pitch and got something else

    Microsoft could have written a clean speed story. They have the numbers for it. Instead the framing is that intelligent apps need a new shape of work, and that shape is built by humans who lead differently, not by humans who type faster.

    That is not a marketing flourish. That is the part of the message that decides whether your AI rollout pays off or quietly burns budget. I have been writing about decision ownership for a while now, and this post is the closest I have seen Microsoft come to saying it out loud.

    Speed is a trap when the underlying decision rights have not moved

    Here is what I keep running into. A team takes a slow approval process. They drop an agent in the middle of it. The agent now drafts the recommendation in two seconds. Approval still takes four days because three managers still need to sign off, and none of them changed how they review.

    You did not speed up the process. You sped up the part nobody was waiting on.

    Worse, the agent now produces ten times the volume of recommendations the approval chain was sized for. The queue grows. People rubber-stamp to keep up. The quality of the decision drops while the appearance of throughput goes up. I have written before that automating a bad process just makes it fail faster. Intelligent apps make this failure mode worse, not better, because the speed gap between the AI step and the human step gets wider. This dynamic is one reason RPA vs AI automation comparisons often miss the point — neither technology fixes a process where decision rights have not moved.

    If decision rights do not move, speed is a trap.

    What human leadership actually has to do for an intelligent app to work

    The Microsoft post uses the phrase human leadership as if everyone knows what it means. I do not think we do. So here is what I think it has to mean operationally for an intelligent app to actually pay off.

    First, someone with authority has to redraw the decision boundary. Which calls does the agent make on its own. Which calls go to a human. Which calls require two humans. That is not a developer task. That is a leadership task, and most orgs avoid it because it is uncomfortable.

    Second, the constraints the agent operates under have to be owned by a person, not buried in a system prompt. This is exactly why Microsoft’s business skills in Dataverse matter. They give policy a home with an owner and a version history. Without that, your intelligent app is running on tribal knowledge that nobody can update.

    Third, leaders have to stop measuring the team on volume of approvals or tickets closed. If the agent is doing the routine work, the human metric has to shift to quality of exception handling and quality of policy. Otherwise you are paying senior people to do work an agent already did.

    None of this is a feature you ship. All of it is org design. The Power Platform tooling will not do it for you.

    The teams I see getting this right are doing one specific thing first

    The teams I talk to who are actually getting value out of intelligent apps and human leadership do one thing before they build anything. They write down, on one page, who currently owns each decision in the process and who will own it after the agent ships. Same column, different rows. The delta is the work.

    That page is uncomfortable to produce. It surfaces the fact that some managers are about to lose a piece of their job, that some policies have no clear owner, and that some approval steps exist only because nobody ever questioned them. This is the part most teams skip, because it is political, not technical. It is also why most Power Platform Center of Excellence setups stall in month three — the governance conversation requires the same political work that most teams defer until it is too late.

    The teams that skip it ship a working app and wonder six months later why nothing changed. The teams that do it ship a smaller app and quietly reshape how a department works.

    So my read on the Microsoft post is this. They did not bury the lede by accident. The lede is human leadership. The intelligent app is the part of the story everyone is comfortable talking about. The other half is the part that decides whether any of this matters.

    If you want to see how I think about this kind of org-shaped problem, more of my writing is on LinkedIn. The technology is rarely the bottleneck. The willingness to move decisions is.

    Frequently Asked Questions

    What is the relationship between intelligent apps and human leadership in the workplace?

    Intelligent apps only deliver real value when human leadership changes how decisions are made, not just how fast tasks are completed. Without redesigning decision rights and workflows, AI tools tend to accelerate broken processes rather than fix them.

    Why does automating an existing process sometimes make things worse?

    Automation increases the speed and volume of outputs, but if the approval or review process stays the same, the bottleneck simply gets worse. Teams end up rubber-stamping decisions to keep pace, which lowers quality while creating the illusion of better throughput.

    How do I know if my organisation is ready for an AI automation rollout?

    A good starting point is asking whether decision rights have been clearly assigned and whether leaders are willing to redesign how approvals and reviews work. If those structures are unchanged, adding AI tools is likely to surface existing problems faster rather than solve them.

    When should I redesign workflows before deploying intelligent apps?

    Workflow redesign should happen before deployment, not after. If the human steps in a process have not been updated to match the speed and volume AI generates, the technology will outpace the people it is meant to support.

    This post was inspired by Intelligent apps, human leadership, and the new shape of work via Microsoft Power Platform Blog.

  • RPA vs AI Automation for Enterprise Workflows

    RPA vs AI Automation for Enterprise Workflows

    RPA vs AI automation comparison for enterprise workflows

    The decision I keep watching teams get wrong: should this workflow be built with RPA or with an AI agent. The RPA vs AI automation debate gets framed as old tech versus new tech, which is the wrong frame entirely. They solve different problems. Picking the wrong one is how you end up with a fragile bot that needs babysitting or an agent that hallucinates its way through invoice approvals.

    I have built both inside a large org. Here is how I actually decide.

    Determinism and predictability

    RPA assumes the screen, the field, and the click path are the same every time. If the SAP transaction code is VA01 today and VA01 tomorrow, RPA wins. It will execute that path 10,000 times with zero variance.

    AI automation assumes variance is the input. The email phrasing changes, the PDF layout changes, the customer asks the same thing five different ways. An agent reasons over that variance. It is non-deterministic by design, which is a feature for unstructured input and a liability for structured execution.

    Rule of thumb I use: if I can write the decision tree on a whiteboard in 15 minutes, it is RPA work. If the decision tree has more than 30 branches and half of them are “it depends on the wording,” it is agent work.

    Cost per execution

    Dimension RPA (Power Automate Desktop) AI Agent (Copilot Studio)
    Per-run cost Near zero after license Roughly 1 message credit per turn, often 5 to 15 turns per task
    License model Per-bot or per-user attended/unattended Message packs, 25,000 messages per pack
    Scaling cost Linear with bot count Linear with conversation volume and tool calls
    Failure cost Bot stops, you fix it Agent confidently completes the wrong task

    RPA at 100,000 runs a month is basically free compute after the license. An agent at 100,000 runs is not. I have seen teams underestimate this by an order of magnitude because they tested with 50 runs and extrapolated linearly without counting tool calls and orchestration turns.

    Maintenance and brittleness

    RPA breaks when the UI changes. A vendor pushes a new SAP Fiori update, three selectors shift, your bot fails at 3am. I have lived this. The fix is usually 30 minutes, but you need someone on call who knows the bot.

    AI agents break differently. They do not fail loudly. They drift. The model provider updates, your prompt that worked last month now produces a slightly different output format, and downstream parsing silently fails. I wrote about this in my agentic workflow post. The failure mode is worse because users find out three days later when the wrong invoice gets paid. If you are building flows that sit underneath an agent, Power Automate error handling patterns that actually work will save you from the silent failures that surface weeks after go-live.

    RPA maintenance is reactive and obvious. Agent maintenance is proactive and requires evaluation infrastructure most teams do not build.

    What the work actually looks like

    This is the dimension nobody compares on. Look at the input.

    Structured input, structured output, no judgment needed: RPA. Copying 200 rows from a legacy system into a SharePoint list, kicking off a daily report, screen-scraping a vendor portal that has no API. Boring, repetitive, deterministic. Power Automate Desktop handles this all day. If you are still deciding whether to invest time in the broader platform, RPA is not the right tool for every repetitive task is worth reading before you commit to a build.

    Unstructured input, structured output, judgment needed: AI. Reading 500 supplier emails and extracting the PO number, classifying tickets by intent, summarizing a 40-page contract into five bullet points. This is where Copilot Studio or a custom agent earns its cost.

    The hybrid case is the most common one and the one most teams miss. The agent reads the email, extracts the structured fields, then hands off to an RPA bot or a cloud flow that executes the deterministic part. The agent is the reasoning layer. RPA is the execution layer. They are not competitors. They are stacked.

    Governance and auditability

    RPA logs are simple. Action ran, action succeeded, here is the screenshot. Auditors love this.

    AI agents need decision logs, not just execution logs. You need to capture why the agent picked tool A over tool B. Most teams I talk to are not logging this and will get caught when the first compliance review hits. I covered this in The Real Shift Is Not Faster Work It Is Who Owns the Decision. Based on what I have built, this is the gap that bites you 6 months in, not on day one.

    Choose RPA if / Choose AI if

    Choose RPA if: the input is structured, the path is deterministic, the volume is high, the cost per run needs to be near zero, and the system has no API. This is most legacy integration work.

    Choose AI automation if: the input is unstructured, the work requires classification or extraction or summarization, variance is the norm, and you have the evaluation discipline to catch silent drift.

    Choose both if: you have a real workflow. Most enterprise automation is hybrid. The line is not RPA versus AI. It is figuring out which layer does what.

    Frequently Asked Questions

    What is the difference between RPA vs AI automation for enterprise workflows?

    RPA is built for repetitive, predictable tasks where the process follows the same steps every time, while AI automation handles unstructured or variable inputs that require reasoning. They are not competing technologies but tools suited to different problems. Choosing the wrong one leads to either a fragile bot or an agent making confident mistakes.

    When should I use RPA instead of an AI agent?

    Use RPA when your process is consistent, rule-based, and can be mapped out as a clear decision tree. If the same fields, screens, or steps repeat thousands of times without variation, RPA will be faster, cheaper, and more reliable than an AI agent.

    How do I know if AI automation is worth the cost for my workflow?

    AI agents consume message credits per turn and most tasks require multiple turns, so costs scale quickly at high volumes. Before committing, calculate expected monthly runs and multiply by average turns per task, not just per conversation. Teams often underestimate this significantly when testing at small scale.

    Why does RPA break so often in enterprise environments?

    RPA relies on fixed UI selectors, so any interface update from a vendor can shift elements and cause the bot to fail. These failures are usually quick to fix but require someone familiar with the bot to be available when issues occur. Unlike AI agents, RPA fails loudly and immediately rather than silently producing wrong results.

  • Claude vs ChatGPT Is the Wrong Question When You Are Building Automations

    Claude vs ChatGPT Is the Wrong Question When You Are Building Automations

    Comparing Claude vs ChatGPT for automation workflows inside Power Platform

    Another Claude vs ChatGPT comparison landed in my feed this week. I came across a piece on the Zapier Blog running the usual head-to-head: reasoning, coding, writing, ethical dilemmas. Useful if you are picking a chat assistant for personal use. Almost useless if you are deciding claude vs chatgpt for automation inside a real enterprise flow.

    I keep seeing people pick a model based on a consumer benchmark and then act confused when their Copilot Studio agent starts returning malformed JSON in week three. The criteria that matter when a model sits behind a connector are not the criteria that make for a good blog post.

    Why Head to Head Model Comparisons Stop Being Useful the Moment You Add a Connector

    Consumer comparisons test the model in isolation. One prompt in, one answer out, a human judges the output. That setup tells you nothing about what happens when the model has to call a tool, parse a response, call another tool, and feed a structured result into a downstream action.

    Inside an automation, the model is not the product. The model is one component in a pipeline. The question is not which one writes better poetry. The question is which one fails in ways your orchestration layer can actually handle.

    I wrote about this angle in a previous post on agentic workflows. The LLM is the reasoning layer, not the agent. Picking the reasoning layer on vibes from a consumer benchmark is how you end up with a beautifully worded confident response for a task that never completed.

    The Four Things That Actually Matter When a Model Sits Inside an Automation

    These are what I actually test for. None of them show up in head-to-head comparisons.

    Structured output stability under load. Ask the same model for the same JSON schema a hundred times with slightly different inputs. Count how often it adds a trailing comma, drops a required field, wraps the JSON in a code fence, or decides today is the day to add a helpful explanation before the response. This is the single biggest source of silent failures I see in production.

    Tool-calling predictability with multiple connectors. Give the model five tools. Watch how it picks. A model that is 95 percent accurate on tool selection with two tools can drop to 70 percent with five because the descriptions start competing. Consumer tests never measure this.

    Behaviour when context gets long. Most real flows accumulate context: user input, previous tool results, system instructions, retrieved documents. I want to know how the model behaves at 40k tokens of accumulated state, not at 500. Instruction drift usually shows up here first.

    Pricing behaviour under loops. An agent that retries three times on a failed tool call can quietly 10x your cost. The cheaper model on paper is not always the cheaper model in production once you account for retry patterns and token accumulation. Latency Is the Quiet Killer of Agentic Workflows covers how round-trip costs compound in ways most people never budget for until it is too late.

    How I Pick Between Claude and GPT for a Specific Flow

    I do not pick a model for the whole platform. I pick per use case.

    For long-context reasoning where the model needs to hold a lot of state and follow detailed instructions without drifting, Claude has been the more predictable option in my testing. Fewer surprise deviations from the system prompt when the context gets messy. If you want to go deeper on why Claude works well as a reasoning layer inside enterprise pipelines, Claude as an Orchestration Brain Is the Most Interesting Thing Happening in Enterprise AI Right Now gets into the architecture side of that decision.

    For fast, cheap, high-volume classification or extraction where the schema is simple and the input is short, GPT models tend to win on cost-per-call and latency. If the task is “read this email and return one of five categories,” I am not paying for a heavyweight reasoning model.

    For tool-calling inside a Copilot Studio agent with multiple Power Automate actions, I test both. There is no universal winner. It depends on how the tool descriptions are written, how many there are, and how ambiguous the user input gets.

    The honest answer most of the time is: it does not matter as much as the people arguing about it think it does. The bigger wins come from tool design, prompt structure, and failure handling. A well-designed flow with a mid-tier model beats a sloppy flow with the flagship every time.

    What to Test Before You Commit a Model to Production

    Before a model goes behind a production flow, I run four checks. Not benchmarks. Checks against the actual flow.

    Run the real schema a hundred times with production-like inputs. Measure malformed output rate. Anything above one percent and you need a validation and retry layer, no matter which model you picked.

    Run the tool-calling logic with the real connector set, not a simplified test set. Watch for the model picking the wrong tool when two descriptions overlap. This is where I lost the most time the hard way.

    Simulate a long session. Feed it accumulated context that looks like a real user journey, not a single clean turn. Watch for instruction drift.

    Load test with the pricing model in mind. Know what a retry storm costs you before it happens in production, not after finance asks questions. The Power Automate documentation covers retry policies, but most people never configure them until something breaks.

    The Claude vs ChatGPT question is the wrong frame. The right question is: which model handles the specific shape of failure my flow is most exposed to. Answer that and the comparison stops mattering. That is the part I keep trying to explain when people ask me, and it still gets pushed aside for whichever model topped a benchmark last week.

    Frequently Asked Questions

    Which is better, Claude vs ChatGPT for automation workflows?

    Choosing between Claude and ChatGPT for automation is less about which model performs better in general benchmarks and more about how each behaves inside a pipeline. The criteria that matter are structured output reliability, tool-calling accuracy, and how well the model holds instructions as context grows. Testing both models against your specific workflow conditions will tell you far more than any consumer comparison.

    Why does my AI agent start producing errors after working fine at first?

    This often happens because the model experiences instruction drift as context accumulates over time. Long flows gather user inputs, tool results, and retrieved documents, and some models struggle to maintain consistent behaviour at high token counts. Testing your model under realistic context lengths before going to production can help catch this early.

    How do I choose an AI model for a Power Automate or Copilot Studio flow?

    Focus on how the model handles structured outputs, selects the right tools when multiple connectors are available, and behaves when context is long rather than short. Consumer benchmarks test models in isolation, but real automation pipelines require consistent, predictable behaviour across repeated calls with varying inputs. Running your own tests against your actual schema and tools will give you more reliable answers.

    What causes silent failures in AI automation workflows?

    One of the most common causes is inconsistent structured output, where a model occasionally adds unexpected formatting, drops required fields, or wraps a response in a code block instead of returning clean JSON. These errors can pass through without triggering obvious alerts while still breaking downstream actions. Testing output stability across many varied inputs is one of the most important steps before deploying a model-powered flow.

    This post was inspired by Claude vs. ChatGPT: What’s the difference? [2026] via Zapier Blog.