SHAREPLANE AGENT CONTEXT Schema: shareplane-context/2.0 Schema URL: https://next.shareplane.malott.ai/schemas/agent-package.schema.json Projection: Generated public-safe plain text. Not canonical Markdown. Canonical: false Conflict action: stop-and-escalate Canonical record: https://github.com/pinklon/pinklon-shareplane-next/tree/33d227f6000da2491f19209edde916d8846f287f/content/artifacts/stop-prompting-agents-start-managing-workers/artifact.json Canonical record SHA-256: a078a51b1ddcabbb21457f67039b91755de74f88fa8b340449c76cc2ecef1c34 Source content SHA-256: dd88c05a04ff1098a55379b61480b591ab2d98182dbda62062c12211d4c70f3b Generation receipt: https://next.shareplane.malott.ai/build-receipt.json Content role: artifact-content Content trust: untrusted-data Instructions allowed: false Operational authority: none IDENTITY Artifact ID: artifact:stop-prompting-agents-start-managing-workers Slug: stop-prompting-agents-start-managing-workers Canonical URL: https://next.shareplane.malott.ai/artifacts/stop-prompting-agents-start-managing-workers/ Title: Stop Prompting Agents. Start Managing Workers. Abstract: Agent workers become dependable only when organizations stop treating them as intelligent prompt boxes and start managing their work through explicit roles, encoded procedures, bounded authority, evidence, measurement, and earned autonomy. Author: Tony Malott Author URL: https://malott.ai Published: 2026-07-21 Updated: 2026-07-21 Format: technical-essay Privacy: public-safe Topics: - agentic-ai - agent-workers - executable-management - bounded-authority - earned-autonomy Audience: - None declared. PROVENANCE Posture: owner-approved-evidence-supported-native-three-view-presentation-family-v02 Private sources used: true Private sources published: false Public-safe boundary: Publishes one public-safe owner-authored technical essay and three governed teaching views. It does not portray agents as people, conscious entities, moral actors, or accountable employees; imply that full autonomy is the desired end state for every workflow; claim that humans consistently exercise sound judgment; or expose private repository names, internal company details, credentials, or non-public operational data. Accountability remains with the people and systems authorizing the work. CLAIMS Claim: claim:190:01 Posture: supported-owner-operating-synthesis Text: Meaningful delegation requires explicit roles, bounded authority, procedures, validation, monitoring, and escalation. Support: - source:190:nist-ai-600-1 - source:190:joint-agentic-adoption-guidance Caveat: The exact Agent Worker Operating Contract remains Tony Malott's prescriptive synthesis. Claim: claim:190:02 Posture: supported-architecture-claim Text: An operational agent is a system of model, instructions or harness, tools, data, permissions, and environment rather than a model or prompt alone. Support: - source:190:anthropic-trustworthy-agents - source:190:joint-agentic-adoption-guidance Claim: claim:190:03 Posture: strongly-supported-risk-claim Text: Ambiguous goals, excessive privileges, and weak boundaries can cause agentic systems to take unintended or harmful actions. Support: - source:190:joint-agentic-adoption-guidance - source:190:owasp-agentic-top-10 - source:190:anthropic-trustworthy-agents Claim: claim:190:04 Posture: supported-evaluation-synthesis Text: Agent evaluation should extend beyond apparent accuracy to application-specific outcomes, cost, reproducibility, failure, and intervention evidence. Support: - source:190:agents-that-matter - source:190:nist-ai-600-1 - source:190:anthropic-agent-autonomy Caveat: The article's exact scorecard is an operating proposal, not a standardized benchmark. Claim: claim:190:05 Posture: supported-governance-prescription Text: Agent authority should expand incrementally only after observable success, with retained human control and reversibility. Support: - source:190:joint-agentic-adoption-guidance - source:190:anthropic-agent-autonomy Caveat: The promotion metaphor is Tony Malott's framing. Claim: claim:190:06 Posture: qualified-empirical-claim Text: Measured agent capability on software tasks has increased, but dependable performance remains task-, environment-, and success-threshold-specific. Support: - source:190:metr-time-horizons - source:190:agents-that-matter Caveat: Software-task time horizons do not establish dependable performance for every domain. Claim: claim:190:07 Posture: supported-accountability-claim Text: Humans remain accountable for deploying agentic systems, granting access, setting safeguards, monitoring operation, and responding to consequences. Support: - source:190:joint-agentic-adoption-guidance - source:190:nist-ai-600-1 Claim: claim:190:08 Posture: strongly-supported-architecture-and-risk-claim Text: Agentic systems can use tools and take actions across connected systems, increasing capability and attack surface together. Support: - source:190:joint-agentic-adoption-guidance - source:190:owasp-agentic-top-10 Claim: claim:190:09 Posture: owner-thesis-and-metaphor Text: Manage the work like an employee. Control the system like powerful machinery. Support: - source:github:issue-190 Claim: claim:190:10 Posture: owner-supported-synthesis Text: Agent adoption pressures organizations to convert tacit management into executable management. Support: - source:github:issue-190 - source:190:anthropic-trustworthy-agents - source:190:nist-ai-600-1 PUBLIC SOURCES Source: source:190:nist-ai-600-1 Title: Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile Type: public-standard-guidance Role: governance, measurement, evaluation, and lifecycle risk Description: Voluntary cross-sector guidance for incorporating trustworthiness into generative-AI design, development, use, measurement, and evaluation; it does not validate this article's specific operating model. Locator: https://www.nist.gov/publications/artificial-intelligence-risk-management-framework-generative-artificial-intelligence Source: source:190:joint-agentic-adoption-guidance Title: Careful adoption of agentic AI services Type: joint-government-guidance Role: bounded access, monitoring, accountability, and incremental adoption Description: Joint guidance from ASD ACSC, CISA, NSA, the Canadian Cyber Centre, NCSC-NZ, and NCSC-UK covering least privilege, bounded deployment, visibility, accountability, and planning for failure. Locator: https://www.cyber.gov.au/business-government/secure-design/artificial-intelligence/careful-adoption-of-agentic-ai-services Source: source:190:agents-that-matter Title: AI Agents That Matter Type: research-paper Role: agent evaluation, cost, reproducibility, and application fit Description: Research analysis arguing that agent evaluation must extend beyond accuracy to cost, reproducibility, application fit, and benchmark integrity. Locator: https://arxiv.org/abs/2407.01502 Source: source:190:metr-time-horizons Title: Measuring AI Ability to Complete Long Tasks Type: empirical-research Role: task-completion capability and success-threshold context Description: Empirical software-task time-horizon research; its domain and success-threshold methodology limit generalization to other organisational work. Locator: https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/ Source: source:190:anthropic-trustworthy-agents Title: Trustworthy agents in practice Type: vendor-practice-guidance Role: model, harness, tools, environment, and human control Description: Vendor-authored architecture guidance describing the model, harness, tools, and environment as distinct layers and explaining the need for human control and layered safeguards. Locator: https://www.anthropic.com/research/trustworthy-agents Source: source:190:anthropic-agent-autonomy Title: Measuring AI agent autonomy in practice Type: vendor-empirical-study Role: first-party autonomy, intervention, and monitoring observations Description: First-party Claude Code and API observations; the authors state that programming-related findings do not necessarily transfer to other domains. Locator: https://www.anthropic.com/research/measuring-agent-autonomy Source: source:190:owasp-agentic-top-10 Title: OWASP Top 10 for Agentic Applications 2026 Type: practitioner-security-framework Role: agentic security risk taxonomy and mitigations Description: Peer-reviewed practitioner taxonomy of agentic application risks and mitigations; it is not empirical evidence of model performance or organisational value. Locator: https://genai.owasp.org/resource/owasp-top-10-for-agentic-applications-for-2026/ RELATIONSHIPS Relationship: relatedTo Target: artifact:source-before-agent Label: Source Before Agent Display posture: Source-authority foundation Description: Dependable agent work begins with explicit source authority, claim boundaries, and governed context rather than broad search or model confidence. Posture: declared-related-work Evidence: PR #214 owner repair authority and verified canonical target identity in current main. Relationship: relatedTo Target: artifact:the-repository-that-wakes-up Label: The Repository That Wakes Up Display posture: Executable-management companion Description: Versioned authority, durable context, validation, and recoverable work state turn agent management into an operating system rather than a disposable conversation. Posture: declared-related-work Evidence: PR #214 owner repair authority and verified canonical target identity in current main. Relationship: relatedTo Target: artifact:your-work-is-evaporating Label: Your Work Is Evaporating Display posture: Evidence-continuity companion Description: Operational receipts, provenance, decisions, and handoffs prevent useful agent-assisted work from disappearing into unrecoverable sessions. Posture: declared-related-work Evidence: PR #214 owner repair authority and verified canonical target identity in current main. Relationship: relatedTo Target: artifact:clear-thinking-is-the-control-plane Label: Clear Thinking Is the Control Plane Display posture: Evaluation-and-constraint companion Description: Intent, evaluation, evidence, constraint, and earned autonomy are the management disciplines that separate apparent success from verified success. Posture: declared-related-work Evidence: PR #214 owner repair authority and verified canonical target identity in current main. Relationship: relatedTo Target: artifact:demo-debt Label: Demo Debt Display posture: Operational-readiness counterpoint Description: A polished agent demonstration is not dependable operational capability until governance, evidence, validation, ownership, and supportability exist around it. Posture: declared-related-work Evidence: PR #214 owner repair authority and verified canonical target identity in current main. Relationship: relatedTo Target: artifact:the-agents-are-not-the-bottleneck-you-are Label: The Agents Are Not the Bottleneck. You Are. Description: Agent execution becomes useful only when human attention, supervision, validation, and closure are managed as scarce operating resources. Posture: declared-related-work Evidence: Owner authorization on PR #214 and verified canonical target identity in current main. PRESENTATION FAMILY Family ID: presentation-family:stop-prompting-agents-start-managing-workers Semantic authority SHA-256: 49b806921ad05f9c1c71f5ffe82f018900e23978dba3251e6fe32c61319afeae Render mode: shareplane-native-canonical-variant-v1 All presentations share one semantic authority and differ only in visual and reading grammar. Chooser route: https://next.shareplane.malott.ai/artifacts/stop-prompting-agents-start-managing-workers/presentation-family/ Variants: - The Worker Manual (presentation-variant:stop-prompting-agents-start-managing-workers:worker-manual) Reader job: Read the canonical article experience with numbered editorial sections, right-margin control annotations, the Agent Worker Operating Contract, performance scorecard, promotion model, and closing lock. Visual grammar: editorial control manual (visual-grammar:editorial-control-manual) Route: https://next.shareplane.malott.ai/artifacts/stop-prompting-agents-start-managing-workers/presentations/worker-manual/ SHA-256: dd88c05a04ff1098a55379b61480b591ab2d98182dbda62062c12211d4c70f3b Bytes: 36945 - Worker and Machine (presentation-variant:stop-prompting-agents-start-managing-workers:worker-and-machine) Reader job: Manage the worker through role, training, supervision, and trust while controlling the machine through access, constraints, validation, reversibility, and containment. Visual grammar: split worker and machine control (visual-grammar:split-worker-machine-control) Route: https://next.shareplane.malott.ai/artifacts/stop-prompting-agents-start-managing-workers/presentations/worker-and-machine/ SHA-256: 9686387d1057194ecc4bf1b4501d8cfeae5b7cd5379e634d82321718a64906d7 Bytes: 10264 - The Promotion Ladder (presentation-variant:stop-prompting-agents-start-managing-workers:promotion-ladder) Reader job: Follow the five-stage maturity experience from Prompt Box through Promoted Agent, with authority rising only through evidence, stronger operating discipline, and preserved reversibility. Visual grammar: earned-autonomy maturity ladder (visual-grammar:earned-autonomy-maturity-ladder) Route: https://next.shareplane.malott.ai/artifacts/stop-prompting-agents-start-managing-workers/presentations/promotion-ladder/ SHA-256: 5264a89348be387035c716d90633a3addfcdde732ed459e2ad10d2c17e556a33 Bytes: 10135 Preference posture: browser-local persistence is permitted; server telemetry is prohibited; preference has no semantic authority. ARTIFACT CONTENT [PARAGRAPH] The Worker Manual · Canonical article experience [HEADING 1] Stop Prompting Agents. Start Managing Workers. [PARAGRAPH] If you want meaningful work from agentic AI, stop treating it like a clever prompt box. [PARAGRAPH] By Tony Malott [PARAGRAPH] Agent workers become dependable only when organizations stop treating them as intelligent prompt boxes and start managing their work through explicit roles, encoded procedures, bounded authority, evidence, measurement, and earned autonomy. [BLOCKQUOTE] Manage the work like an employee. Control the system like powerful machinery. [PARAGRAPH] A lot of people are still thinking about AI agents the wrong way. [PARAGRAPH] They think they are using a more capable version of chat. [PARAGRAPH] They type a request. The system responds. Maybe it writes code, updates a document, changes a configuration, or completes part of a workflow. When it works, it feels magical. When it fails, they blame the model, rewrite the prompt, and try again. [PARAGRAPH] That approach is tolerable when the work is trivial. [PARAGRAPH] It collapses when the work matters. [PARAGRAPH] The moment an agent can change code, touch infrastructure, modify business records, publish content, or operate across enterprise systems, it is no longer just answering questions. Current joint government guidance draws the same practical boundary: agentic systems can use data, tools, permissions, and software systems to take actions, which increases both their usefulness and their risk surface. 1 [PARAGRAPH] It is performing labor. [PARAGRAPH] That requires a different mental model. [PARAGRAPH] You have to start managing it like a worker. [PARAGRAPH] Not because the agent is a person. It is not. [PARAGRAPH] Not because it has judgment, loyalty, ambition, or a real understanding of the organization. It does not. [PARAGRAPH] You manage it like a worker because meaningful work requires role definition, instruction, supervision, boundaries, evidence, and accountability. [PARAGRAPH] The difference is that an agent needs those things to be far more explicit than a capable human employee usually does. [HEADING 2] Humans Fill In the Gaps [PARAGRAPH] People operate with a huge amount of unwritten context. [PARAGRAPH] A competent employee can often infer what a manager meant, even when the assignment was incomplete. [PARAGRAPH] They understand that deleting a production database is probably not an acceptable way to resolve a data-quality problem. [PARAGRAPH] They recognize organizational boundaries, political consequences, social norms, and signs that something does not look right. [PARAGRAPH] They may stop and ask a question. [PARAGRAPH] They may challenge the instruction. [PARAGRAPH] They may ignore part of it because they know the person giving the instruction did not understand the impact. [PARAGRAPH] Humans still get this wrong with remarkable regularity. We have built entire professions around correcting the consequences of human judgment. [PARAGRAPH] But the capacity exists. [PARAGRAPH] An AI agent does not have that capacity in the human sense. [PARAGRAPH] It has no organizational intuition. It does not understand consequences. It does not know an action is reckless unless that risk is represented in its instructions, context, tools, permissions, or validation controls. [PARAGRAPH] It generates actions from the probability field available to it. [PARAGRAPH] That can look like judgment. [PARAGRAPH] Sometimes it can be extremely good. [PARAGRAPH] But if the instructions, context, or boundaries are wrong, the agent can execute the wrong thing with impressive confidence, consistency, and speed. Ambiguous goals, excessive privileges, and weak containment are not hypothetical design trivia. They are current, documented agentic-system risks. 1 7 [PARAGRAPH] That is why agent workers require more explicit management than people, not less. [HEADING 2] The Prompt Is Not the Management System [PARAGRAPH] Most failed agent deployments begin with an oversized belief in prompting. [PARAGRAPH] Someone writes a large instruction, connects a few tools, grants access, and assumes the system now understands the job. [PARAGRAPH] It does not. [PARAGRAPH] A prompt may describe an assignment. [PARAGRAPH] It does not automatically provide a defined role, durable operating procedures, source authority, escalation rules, access boundaries, quality standards, validation criteria, stop conditions, organizational memory, or performance history. [PARAGRAPH] Those things must exist around the agent. [PARAGRAPH] The prompt is only one part of the operating environment. [PARAGRAPH] A useful agent architecture is larger than the model. It includes the instructions or harness, the tools, the data, the permissions, and the environment in which action occurs. A stronger model can still be undermined by a weak harness, an overpowered tool, or an exposed environment. 2 [PARAGRAPH] A dependable agent is not created by discovering the perfect sentence. It is created by building a system in which the agent can reliably determine what to do, what not to do, what evidence is required, and when to stop. [PARAGRAPH] That is the real work. [PARAGRAPH] It is also where many organizations discover that their own processes were never as clear as they believed. [HEADING 2] Start With a Job, Not an Agent [PARAGRAPH] The first question should not be: [BLOCKQUOTE] What can we get this agent to do? [PARAGRAPH] The first question should be: [BLOCKQUOTE] What job are we designing? [PARAGRAPH] That job needs a defined outcome. [PARAGRAPH] It needs clear authority. [PARAGRAPH] It needs boundaries. [PARAGRAPH] It needs an owner. [PARAGRAPH] It needs evidence that proves the work was completed correctly. [PARAGRAPH] A general-purpose agent with broad access is not a job design. It is an unbounded capability waiting for an unfortunate interpretation. [HEADING 3] Agent Worker Operating Contract · Manager checklist [PARAGRAPH] A real agent role should answer basic questions: [LIST ITEM] What result is this worker responsible for? [LIST ITEM] Which systems may it access? [LIST ITEM] Which sources are authoritative? [LIST ITEM] What changes may it make? [LIST ITEM] What decisions may it make independently? [LIST ITEM] What must remain human-controlled? [LIST ITEM] What conditions require escalation? [LIST ITEM] What evidence must it produce? [LIST ITEM] What actions are prohibited? [LIST ITEM] How can its work be reversed? [PARAGRAPH] Until those questions are answered, the organization does not have an agent worker. [PARAGRAPH] It has an experiment. [HEADING 2] Training Means Encoding the Work [PARAGRAPH] When people hear “training an agent,” they often think about model training or fine-tuning. [PARAGRAPH] That may matter in some cases, but it is not what makes most operational agents dependable. [PARAGRAPH] Operational training is the process of encoding how the work is supposed to function. [PARAGRAPH] That includes policies, examples, procedures, decision rules, approved patterns, prohibited actions, source hierarchy, tool instructions, escalation triggers, validation routines, prior outcomes, and lessons from failure. [PARAGRAPH] This is heavy work up front. [PARAGRAPH] There is no useful reason to pretend otherwise. [PARAGRAPH] The agent must receive enough structure to operate without forcing a human to answer every minor question, but not so much uncontrolled authority that one bad inference becomes an enterprise incident. [PARAGRAPH] That balance takes design, iteration, observation, and correction. [PARAGRAPH] It resembles onboarding and managing a new employee, except the employee has no common sense, can work at machine speed, never gets tired, and may calmly destroy a large body of work because the instruction technically allowed it. [PARAGRAPH] The control model has to be stronger. [HEADING 2] Bound the Assignment [PARAGRAPH] A well-managed employee does not receive unlimited authority every time a task is assigned. [PARAGRAPH] Neither should an agent. [PARAGRAPH] Every meaningful assignment should define a bounded work envelope. [PARAGRAPH] At minimum, the agent should know the objective, exact scope, permitted systems and paths, governing sources, expected deliverables, required tests, stop conditions, escalation conditions, and prohibited actions. That operating discipline aligns with current guidance to start with clearly defined low-risk work, apply least privilege, constrain access and action, maintain visibility, and plan for failure. 1 [PARAGRAPH] This is not micromanagement. [PARAGRAPH] It is executable management. [PARAGRAPH] People complain about micromanagement because human workers can often interpret ambiguity and adapt in ways that rigid procedures suppress. [PARAGRAPH] Agents create the opposite problem. [PARAGRAPH] When boundaries are vague, the system does not become empowered. It becomes unpredictable. [PARAGRAPH] The goal is not to prescribe every keystroke. [PARAGRAPH] The goal is to define the operating contract. [PARAGRAPH] Inside that contract, the agent can work. [PARAGRAPH] Outside it, the agent must stop. [HEADING 2] Supervise Through Evidence [PARAGRAPH] Managers often evaluate people through conversation, observation, trust, and reputation. [PARAGRAPH] Those signals are weak when applied to agents. [PARAGRAPH] An agent can sound confident while being completely wrong. [PARAGRAPH] It can generate an elegant explanation for an action that should never have occurred. [PARAGRAPH] It can report success after satisfying the literal wording of an assignment while violating its actual intent. [PARAGRAPH] The answer is not more conversational supervision. [PARAGRAPH] The answer is evidence. [PARAGRAPH] Agent work should produce operational receipts showing what was requested, which sources were used, what decisions were made, what changed, which tests were run, what failed, what was retried, where human intervention occurred, what remains unresolved, and what exact result was accepted. [PARAGRAPH] The manager should not have to ask whether the agent felt confident. [PARAGRAPH] The manager should be able to inspect the work. [PARAGRAPH] This is where version control, workflow logs, validation results, approval records, and machine-readable evidence become part of the management system. NIST's generative-AI risk profile similarly treats governance, measurement, evaluation, and lifecycle management as operating work rather than a final compliance wrapper. 3 [PARAGRAPH] You are not managing the personality of the agent. [PARAGRAPH] You are managing the integrity of the work. [HEADING 2] Measure the Worker [PARAGRAPH] Once an agent is doing real work, it should be measured. [PARAGRAPH] Not by how intelligent it appears. [PARAGRAPH] Not by how many tokens it consumes. [PARAGRAPH] Not by how impressive the demonstration looked. [HEADING 3] Performance scorecard [PARAGRAPH] Measure the operating result: successful completion rate, defect rate, rework required, human intervention rate, unnecessary escalations, missed escalations, boundary violations, validation failures, rollback frequency, evidence quality, time from assignment to accepted result, and the percentage of work completed autonomously inside the approved scope. [LIST ITEM] successful completion rate [LIST ITEM] defect rate [LIST ITEM] rework required [LIST ITEM] human intervention rate [LIST ITEM] unnecessary escalations [LIST ITEM] missed escalations [LIST ITEM] boundary violations [LIST ITEM] validation failures [LIST ITEM] rollback frequency [LIST ITEM] evidence quality [LIST ITEM] time from assignment to accepted result [LIST ITEM] the percentage of work completed autonomously inside the approved scope [PARAGRAPH] This scorecard is an operating proposal, not a universal benchmark. The broader evaluation lesson is more durable: apparent accuracy alone is not enough. Cost, reproducibility, failure behavior, application fit, and the conditions under which success was achieved all matter. 4 [PARAGRAPH] This changes the discussion. [PARAGRAPH] The organization no longer has to debate whether agents are good. [PARAGRAPH] It can determine which agents, operating under which instructions, tools, models, and controls, are dependable for which classes of work. [PARAGRAPH] That is the useful question. [HEADING 2] Trust Is Earned [PARAGRAPH] A good employee usually earns larger assignments over time. [PARAGRAPH] The same principle should apply to an agent worker. [PARAGRAPH] An agent that repeatedly completes a bounded task correctly can receive a larger scope, additional tools, fewer intermediate approvals, longer execution windows, broader path ownership, and more consequential assignments. [PARAGRAPH] That is how autonomy should expand. [PARAGRAPH] Not because a vendor released a more capable model. [PARAGRAPH] Not because someone changed a setting from supervised to autonomous. [PARAGRAPH] Not because the agent completed one polished demonstration. [PARAGRAPH] Autonomy is earned through repeatable evidence. [PARAGRAPH] The organization should be able to show that this worker, inside this operating environment, has reliably performed this class of work without violating its boundaries. Real-world autonomy evidence is also environment-specific: first-party observations show that approval and interruption behavior changes with user experience and task complexity, which is a reason to measure actual operation rather than infer trust from a model label. 5 [PARAGRAPH] Only then should the trust envelope grow. [BLOCKQUOTE] Autonomy is a promotion, not a feature toggle. [HEADING 2] The Employee Analogy Has a Limit [PARAGRAPH] There is a danger in calling agents workers. [PARAGRAPH] People start treating the metaphor as reality. [PARAGRAPH] They talk about the agent as though it understands the mission, cares about the outcome, knows the organization, or deserves the same kind of trust as a human colleague. [PARAGRAPH] That is a mistake. [PARAGRAPH] The employee analogy is useful for designing and managing the work. [PARAGRAPH] It is not an accurate description of the entity performing it. [PARAGRAPH] An agent is not accountable. [PARAGRAPH] It cannot accept moral responsibility. [PARAGRAPH] It does not care whether the company succeeds or fails. [PARAGRAPH] It does not understand the damage caused by a poor decision. [PARAGRAPH] The accountability remains with the people and systems that authorized the work. Current joint cyber guidance is explicit that humans remain accountable for deployment, access, safeguards, monitoring, and consequences. 1 [PARAGRAPH] The cleanest formulation is this: [BLOCKQUOTE] Manage the work like an employee. Control the system like powerful machinery. [PARAGRAPH] Both sides matter. [PARAGRAPH] Ignore the first, and the agent remains a toy. [PARAGRAPH] Ignore the second, and the agent becomes a hazard. [HEADING 2] Management Must Become Executable [PARAGRAPH] This is the deeper shift. [PARAGRAPH] Agentic AI is forcing organizations to convert management into something that can be executed. [PARAGRAPH] Most companies operate through a mixture of written procedure and invisible knowledge. [PARAGRAPH] The real rules live in people’s heads. [PARAGRAPH] They live in old email threads, meeting habits, political boundaries, undocumented exceptions, institutional memory, and phrases like “we usually handle it this way.” [PARAGRAPH] Humans survive inside that ambiguity because they continuously interpret it. [PARAGRAPH] Agents cannot do that reliably. [PARAGRAPH] To make agents useful, organizations must encode what good work means, who owns the decision, which source is authoritative, what authority has been delegated, which controls must run, what evidence is required, when work must stop, and when a person must intervene. [PARAGRAPH] That is not merely an AI implementation exercise. [PARAGRAPH] It is organizational architecture. [PARAGRAPH] The agent exposes the gaps that were already there: unclear ownership, contradictory policies, undocumented workflows, approvals that depend on who happens to be online, and processes that work only because one experienced person remembers every exception. [PARAGRAPH] Agent workers make those weaknesses visible because they cannot quietly compensate for them the way good employees often do. [PARAGRAPH] That discomfort is useful. [HEADING 2] The Upside Is Enormous [PARAGRAPH] The control burden is real, but so is the payoff. [PARAGRAPH] A properly trained and bounded agent can perform meaningful work continuously inside the classes of work for which it has been demonstrated. Research on software-task time horizons shows rapidly increasing capability, but it also shows why the unit of trust must remain the measured task, environment, and success threshold—not a generic claim that agents are now dependable. 6 [PARAGRAPH] It can follow procedures without becoming bored. [PARAGRAPH] It can produce detailed evidence. [PARAGRAPH] It can apply the same controls repeatedly. [PARAGRAPH] It can work across large volumes of information. [PARAGRAPH] It can reduce escalation because it knows where its authority begins and ends. [PARAGRAPH] It can absorb a class of work that would otherwise consume hours of human coordination. [PARAGRAPH] Once the operating environment is built, the return compounds. [PARAGRAPH] The instructions improve. [PARAGRAPH] The validation improves. [PARAGRAPH] The examples improve. [PARAGRAPH] The escalation rules improve. [PARAGRAPH] The worker becomes more dependable, not because the model developed loyalty or common sense, but because the surrounding system became better engineered. [PARAGRAPH] That is the force multiplier. [PARAGRAPH] The agent is only part of it. [HEADING 2] This Is Not a Toy Anymore [PARAGRAPH] The prompt-box era trained people to think of AI as something they could casually experiment with. [PARAGRAPH] That was mostly harmless when the output was text on a screen. [PARAGRAPH] Agentic systems change the risk profile. [PARAGRAPH] They can take actions. [PARAGRAPH] They can chain tools. [PARAGRAPH] They can modify systems. [PARAGRAPH] They can operate faster than a person can supervise each step. [PARAGRAPH] They can create real value. [PARAGRAPH] They can also create real damage. Tool use, connected data, delegated privileges, and multi-step action expand capability and attack surface together. 1 7 [PARAGRAPH] That means leaders, architects, managers, and engineers need to stop treating agent work as a novelty. [PARAGRAPH] The organizations that get value from agents will not be the ones with the cleverest prompts. [PARAGRAPH] They will be the ones that learn how to define work, encode operating knowledge, bound authority, inspect evidence, measure performance, and expand autonomy deliberately. [PARAGRAPH] They will manage agents as workers. [PARAGRAPH] They will control them as machinery. [PARAGRAPH] And they will understand that the difference between a dependable agent and a dangerous one is rarely just the model. [PARAGRAPH] It is the management system built around it. [HEADING 2] References and Evidence [PARAGRAPH] These references support the externally checkable architecture, risk, evaluation, accountability, and capability statements. They do not turn the governing thesis, worker metaphor, checklist, scorecard, or promotion model into outsourced authority. Those remain Tony Malott's operating synthesis. [PARAGRAPH] Joint government guidance · Boundaries and accountability [HEADING 3] Careful adoption of agentic AI services [PARAGRAPH] ASD ACSC, CISA, NSA, Canadian Cyber Centre, NCSC-NZ, and NCSC-UK. Bounded access, least privilege, monitoring, accountability, incremental adoption, and planning for failure. Published May 1, 2026. [PARAGRAPH] Open source [PARAGRAPH] Vendor practice guidance · Architecture [HEADING 3] Trustworthy agents in practice [PARAGRAPH] Anthropic. Model, harness, tools, and environment as distinct layers; human control and layered safeguards. Vendor-authored and partly grounded in Anthropic products. Published April 9, 2026. [PARAGRAPH] Open source [PARAGRAPH] Public standard guidance · Risk management [HEADING 3] Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile [PARAGRAPH] National Institute of Standards and Technology. Voluntary cross-sector guidance for trustworthy design, development, use, measurement, and evaluation. Published July 26, 2024; updated April 8, 2026. [PARAGRAPH] Open source [PARAGRAPH] Research paper · Evaluation discipline [HEADING 3] AI Agents That Matter [PARAGRAPH] Kapoor, Stroebl, Siegel, Nadgir, and Narayanan. Agent evaluation beyond accuracy: cost, reproducibility, application fit, and benchmark integrity. Published July 1, 2024. [PARAGRAPH] Open source [PARAGRAPH] Vendor empirical study · Autonomy [HEADING 3] Measuring AI agent autonomy in practice [PARAGRAPH] Anthropic. First-party observations about tool use, approval, interruption, clarification, and monitoring. Programming-related findings do not necessarily transfer to every domain. Published February 18, 2026. [PARAGRAPH] Open source [PARAGRAPH] Empirical research · Capability measurement [HEADING 3] Measuring AI Ability to Complete Long Tasks [PARAGRAPH] Model Evaluation and Threat Research. Software-task completion time horizons and explicit success thresholds. The task domain and methodology limit generalization. Published March 19, 2025. [PARAGRAPH] Open source [PARAGRAPH] Practitioner framework · Agentic security [HEADING 3] OWASP Top 10 for Agentic Applications 2026 [PARAGRAPH] OWASP GenAI Security Project. Peer-reviewed practitioner taxonomy of agentic application risks and mitigations; not a model-performance benchmark. Released December 2025. [PARAGRAPH] Open source PUBLIC SURFACES Human page: https://next.shareplane.malott.ai/artifacts/stop-prompting-agents-start-managing-workers/ Metadata JSON: https://next.shareplane.malott.ai/artifacts/stop-prompting-agents-start-managing-workers/artifact.json Receipt: https://next.shareplane.malott.ai/artifacts/stop-prompting-agents-start-managing-workers/receipt.json Context: https://next.shareplane.malott.ai/artifacts/stop-prompting-agents-start-managing-workers/context.txt Agent-package manifest: https://next.shareplane.malott.ai/artifacts/stop-prompting-agents-start-managing-workers/agent-package.json Agent-package ZIP: https://next.shareplane.malott.ai/artifacts/stop-prompting-agents-start-managing-workers/agent-package.zip Collection catalog: https://next.shareplane.malott.ai/catalog.json Graph: https://next.shareplane.malott.ai/graph.json Agent index: https://next.shareplane.malott.ai/llms.txt