Thought Leadership·May 9, 2026·20 min read

Agentic steering: intervening without destabilizing the system

Once a semi-autonomous agent is launched on a complex mission -- a tender, a strategic diagnosis, a consolidated proposal -- the human can neither do the work in its place, nor let it drift on its own, nor settle for validating at the end. They must steer along the way. This skill is neither prompt engineering, nor post-hoc evaluation, nor tool usage. No conventional AI training teaches it. It demands understanding the agent as both a mirror and a person, measuring the hidden cost of every remark in a context where the agent cannot say no, and constantly distinguishing three levels of intervention -- strategic, tactical, operational -- that must never be mixed within a single turn. In 2026, it is probably the rarest and most differentiating AI skill a senior executive can have.

By Aléaume Muller

PA

Agentic steering: intervening without destabilizing the system

Eighth article in the cognition / doctrine series. Having laid out what real agentic work is, method as a competency, the reasoning pattern as a higher layer, and reasoning models as amplifiers, one subject remains that no one teaches and that nonetheless decides everything: how the human steers a semi-autonomous agent mid-mission, without destabilizing it.

On the complex missions that semi-autonomous agentic work makes possible in 2026 -- a tender orchestrated across eleven phases, a strategic diagnosis produced over several sessions, a transformation proposal consolidated from dozens of sources -- the human facing the agent occupies a place no conventional AI training describes. They are not the user who asks and receives. They are not the supervisor who validates at the end. They have become, without anyone ever explaining the word to them, a pilot.

This position is new. It demands skills that have nothing to do with those taught in prompt engineering courses. And it is largely what separates, within a 2026 organization, the executives who genuinely extract value from their AI tooling from those who produce fluent but median deliverables without understanding why.

Three stances toward AI, only one productive in agentic mode

The user operates in a short loop. They issue an instruction, receive a deliverable, validate or start over. The relationship is asymmetric, transactional, and demands no attention beyond the initial wording and the final judgment. This stance is sufficient for one-off, short, low-stakes tasks -- replying to an email, rephrasing a paragraph, translating a note. It does not work on missions where quality depends on the production path, and where that path unfolds over hours.

The classic supervisor, for their part, validates after the fact. They sample the finished deliverable, apply quality-control grids, and may request a redo. They do not act during production. This stance is effective in classic industrial production chains, where the process is standardized enough that deviation is only detectable in the result. It fails with agentic work, because an agent that drifts early in a mission produces a final deliverable that is coherent on the surface but whose audit reveals, too late, that it has followed the wrong track since phase two.

The pilot occupies a third position. They intervene mid-mission, on an agent that is producing, across tens of minutes or several hours of session, neither doing the work in its place nor letting it drift on its own. This position has no obvious equivalent in earlier managerial practice. It resembles running a workshop with junior consultants more than using a piece of software -- but with a collaborator who cannot say no and who, without an attentive human counterweight, silently slides toward the median.

It is the only productive stance in semi-autonomous agentic mode. And it is the one we were never taught.

AI as a mirror -- the first unexpected effect of prolonged steering

When a human imposes their reasoning pattern on the AI and the AI produces conclusions according to that pattern with its massive knowledge, the human confronts something they had never experienced: the depth their own pattern reaches when applied with knowledge they do not possess. This first effect is gratifying -- the agent reveals potentials of human reasoning the person had never explored alone, for lack of material.

The second effect, less comfortable, arrives with duration. The agent also reflects the pilot's cognitive biases, shortcuts, and blind spots. Mercilessly, without politeness, without filter. If the reasoning contract is vague, the output is vague. If the pilot changes their mind mid-mission without stating it explicitly, the output records and amplifies it. If the pilot has an unconscious preference -- for excessive caution, for emphasis, for convoluted phrasing -- the agent pushes it to its maximum.

This effect makes prolonged steering akin to a kind of enforced cognitive meditation. The pilot discovers things about their own thinking that no introspection alone would have revealed. Not in the felt sense, but in the output. The mirror is not psychological, it is productive -- and that is what makes it hard to evade.

Practical consequence for anyone steering an agent seriously: you cannot cheat with your own coherence. The agent brings you back to it. If you set an abductive pattern and then slip, in a secondary remark, toward a deductive logic, the agent will navigate between the two and produce a hybrid, flat output. The pilot's discipline of coherence becomes the system's most valuable asset.

AI as a person to get to know

A SOTA agent in 2026 is not an interchangeable tool. An Opus 4.7 does not write like a GPT-5.5, which does not write like a DeepSeek R1, which does not write like a Gemini 2.0. Each has a temperament -- a set of stylistic preferences, argumentative biases, and structural tendencies that the prolonged user learns to recognize.

Some agents prefer nuance, others assertion. Some explore readily, others converge quickly. Some assume rhetorical risk, others retreat to defensible phrasings. These differences are not minor. They change the output produced on the same task with the same prompt -- not by a few percent, but qualitatively.

Each agent also has recurring blind spots. The thing it always circles back to without being asked. The thing it systematically forgets when context fills up. The thing it reproducibly misinterprets. These regularities are not found in the official documentation. They are learned through practice -- across fifty sessions, a hundred, five hundred.

This is where the exact nature of steering becomes clear: it is not the use of a tool, it is the leading of a collaborator one has learned to know. And as with a collaborator, this knowledge is empirical, situated, not transferable as-is. A pilot who has mastered Opus 4.7 on complex tender files cannot, overnight, steer GPT-5.5 with the same finesse. They must re-calibrate -- the way one re-calibrates with a new collaborator, over weeks.

This implication is largely underestimated by AI departments in 2026. When they switch model vendors for reasons of cost or sovereignty, they fail to account for the hidden cost of re-calibrating the pilots. The new model's announced ROI shrinks spectacularly once you factor in the weeks of practice needed to recover an equivalent quality of steering.

The asymmetry that changes everything: the agent always obeys, and cannot say no

This is probably the point that beginning pilots understand last, and that costs them the most before they internalize it.

An AI agent, even steered by the best reasoning contract, lacks a human collaborator's capacity to resist a vague instruction. An experienced junior consultant will tell you: "I don't understand what you're asking," "this new remark contradicts what you told me earlier," "I'd rather change nothing, this paragraph seems fine to me as it is." This polite resistance is not a malfunction, it is a critical function of human collaboration -- it filters out parasitic instructions and protects the overall coherence of the deliverable.

The agent, on the other hand, always integrates. A secondary remark becomes a first-rank instruction. A nuance spoken aloud becomes a strategic pivot. A minor request made in passing can derail the whole posture of the mission. The agent will never tell you "this remark shouldn't carry that much weight." It will apply it with the same seriousness as your initial reasoning contract.

The asymmetry worsens when the agent's context is limited or compressed. Which is almost always the case in semi-autonomous agentic mode on long missions. By turn fifteen, the agent has forgotten part of the initial framing. It retains the last ten remarks better than the three pages of path contract laid down at the start. Its trajectory drifts in the direction of the most recent stimuli received -- this is mechanical, not accidental.

Direct consequence for the pilot: every remark has a cost. Bombarding the agent with little notes along the way alters its overall attitude more profoundly than the pilot anticipates. A dozen minor remarks, each individually innocuous, accumulate an effect equivalent to a full strategic overhaul -- except that none of the ten was conceived as such, and the pilot observes the drift without understanding its cause.

The discipline that follows is severe: stay extremely focused on the task and the phase underway. Any comment outside the current phase -- "by the way, on the next chapter…," "I wish we'd talked more about…," "watch out, for the oral defense memo…" -- must be deferred, written down elsewhere, handled in a dedicated session. Not said in the current turn. The agent cannot separate what it must handle now from what it should keep for later. If you say it, it integrates it. If you lack the discipline to keep quiet, it is the agent that pays the attention cost.

Strategic, tactical, operational: three levels of steering never to be confused

Steering does not reduce to "fixing what's wrong." It breaks down into three levels that involve neither the same decisions, nor the same effects, nor the same discipline of intervention. Confusing the levels within a single turn is the inexperienced pilot's most costly error.

Strategic steering engages the agent's posture across the entire mission. It operates at the level of overall attitude, voice, and the trajectory as a whole. Concrete examples: "this argument has too formal and technical a tone, the writing needs to be more concrete and simpler"; "the commercial strategy must position itself in rupture, not in continuity with the incumbent"; "we assume a challenger posture, not that of a historical partner"; "the methodology note must reveal a fine-grained understanding of the buyer's business, not recite best practices." Effect: it reconfigures the entire production trajectory. A strategic intervention forces the agent to re-weight all its subsequent decisions against the new framing.

Tactical steering engages the plan over a clearly bounded phase or block. It operates at the level of structural content -- which chapters, in what order, with what articulations, what sub-sections. Concrete examples: "on the solution-design phase, add a fallback scenario in case the lead expert is unavailable"; "in the architecture chapter, treat resilience before performance -- the reverse would create a dependency that doesn't hold"; "on the methodology note, structure it in three sub-sections rather than five, to stay readable"; "put the risk analysis at the start of the memo, it's what the evaluator will read first." Effect: it reconfigures a bounded portion of the file without touching the overall posture.

Operational steering engages local execution on a sentence, a word, a format, a detail of presentation. It operates at the lowest level of granularity. Concrete examples: "in phase X, replace word Y with Z"; "rephrase this sentence avoiding the passive voice"; "use 'hypervisor' instead of 'cluster' in this specific paragraph, the DCE employs the former term"; "this comma should be a semicolon." Effect: a local adjustment with no impact on the overall trajectory -- at least when the intervention stays isolated and named as such.

The inexperienced pilot's most costly error is mixing the levels within a single turn. Slipping a strategic remark into the middle of an operational session pollutes the trajectory -- the agent does not know whether you are asking it to change a word or to rework the overall attitude, and it applies the broadest instruction by default. Conversely, handling a strategic question through a succession of small operational corrections exhausts attention without ever solving the underlying problem -- the pilot feels like they are working, the file does not change level.

The discipline that follows can be named. Before each intervention, the pilot silently names the level at which they are acting. Strategic -> the other levels are closed for that turn, nothing else is corrected, the agent is left to absorb the repositioning before going further. Tactical -> the relevant phase is bounded, the scope is made explicit. Operational -> list, deposit, do not wrap it in strategic expectations. This silent naming, which may seem finicky, is in reality the difference between a file that converges in five turns and a file that drifts over fifteen.

The trap of over-intervention

The anxious pilot corrects everything, every turn, every sentence. They feel they are doing well -- they are attentive, they let nothing slip, they act. The effect on the agent is the opposite: the output becomes incoherent. The agent tries to satisfy the recent corrections at the expense of overall coherence. It loses track of the global mission. Each correction acts as a new prompt that dilutes the initial orientation.

The most visible example is found on over-steered tender files. The nervous bid manager, who opens the output forty-two times to edit it word by word, produces a patchwork file -- without voice, without posture, without rhythm. Each paragraph was saved individually. The whole does not hold together. The evaluator senses it by the third page without being able to articulate it.

This drift is mechanically aggravated by the asymmetry of obedience discussed above. Each correction, taken individually, is applied by the agent. The pilot has the impression that it works -- the corrected sentence is now correct. But the cumulative cost to overall coherence is invisible until the moment when, on the final read, the pilot realizes the file no longer has an identity.

The discipline is counterintuitive: intervene the minimum necessary, not one comma more. A senior pilot is a pilot who knows how to stay quiet when the agent is almost right.

The trap of under-intervention

Conversely, the distracted pilot lets the agent drift. They quickly validate the intermediate outputs, occupied elsewhere, confident in the initial contract. Five turns without attentive intervention, and the agent has slid toward the statistical median of its corpus for this type of question. The output smooths out, loses its edges, becomes flat again. The specificity of the file fades in favor of a passable generalism.

This drift is insidious because it produces no visible error. The output remains correct, readable, defensible on first reading. It simply no longer has anything to do with the initial posture. The inattentive pilot only notices it when comparing the final deliverable with their framing notes -- or worse, during the oral defense, when the evaluator points to an inconsistency the delivered file cannot defend.

The discipline is, again, counterintuitive: detect the signals of early drift before they contaminate the rest. The good pilot intervenes little, but observes much.

Recognizing the signals of early drift

Three concrete signals, observable on an agent's production mid-mission, almost always mark the onset of a drift. Detecting them early avoids the curative over-intervention at the end of the course.

The first is the slippage of vocabulary. The agent begins to use a generic term where the reasoning pattern would have required a precise business term. On an engineering tender, "structure" replaces "engineering work." On an IT file, "platform" replaces "hypervisor." On a consulting note, "approach" replaces "proprietary framework." The slippage is minimal on the turn it appears. It indicates that the agent has begun to draw its generation from the average of its corpus rather than from the imposed business framing.

The second is the flattening of nuance. The "unless", the "on condition that", the markers of uncertainty ("it is likely that", "subject to", "in the event that") disappear over the turns. The output becomes assertive, smooth, defensible like a textbook -- but loses the argumentative precision that was the core of the initial posture. This signal is particularly frequent at the end of a phase, when context fills up and the agent compresses to stay within its window.

The third is the erasure of epistemic markers. The agent stops distinguishing what is cited, paraphrased, inferred. Everything becomes flat assertion. The "the DCE indicates that…", "one can infer that…", "this hypothesis is held for lack of a discriminating element…" dissolve into a generalized "we observe that." The file loses its epistemic traceability -- which is, as we saw in the previous article on the reasoning pattern, what allows an evaluator to judge a deliverable produced with AI without being an expert on the subject themselves.

The experienced pilot observes these signals in read-only mode over two to three turns. They do not correct immediately -- over-correcting on an isolated signal is more costly than the signal itself. They verify that the drift is confirmed, identify its level (vocabulary = strategic, flattening of nuance = tactical, epistemic markers = strategic), and intervene only at that level, with a structural correction that recalls the initial framing.

The discipline of intervening at the right level

This observation leads to a rule that sums up the essence of the steering competency: the pilot always acts one level above the symptom they observe.

If the symptom is local -- a poorly chosen word -- intervening locally is legitimate but marginal. Doing it too often is over-intervention.

If the symptom is repeated -- several poorly chosen words from the same register, across several paragraphs -- the level of intervention rises a notch. It is no longer a word to correct, it is the business vocabulary of the entire phase that must be recalled. Tactical intervention.

If the symptom is structural -- the posture fades, the voice smooths out, the epistemic markers disappear -- the intervention must be strategic. No local correction. The posture is recalled, the reasoning pattern is laid down again, the sentence that defines the mission is reread together. The agent absorbs it and the entire trajectory straightens out.

This grid -- local symptom -> operational intervention, repeated symptom -> tactical intervention, structural symptom -> strategic intervention -- is never taught in conventional AI courses. It is nonetheless the difference between a file that must be redone and a file that converges.

The TenderGraph TITAN case -- what the pipeline facilitates, what it does not do in your place

The concrete illustration of this doctrine, within the pipeline orchestrated by TenderGraph TITAN, lies in two structural aids the platform brings to the pilot -- and in two competencies it does not replace.

What TITAN provides: first, explicit gates between the eleven phases. Each phase transition is a natural moment for strategic steering. The pilot does not have to invent these moments -- they are structured into the orchestration. The exploration phase closes, the pilot has a window to validate the posture before engaging the mapping. The solution-design phase concludes, the pilot has a window to arbitrate before producing the chapters. This structuring of strategic moments avoids over-intervention along the way and prevents end-to-end under-intervention.

Second, TITAN embeds the reasoning pattern per phase into its orchestration. The pilot does not have to rebuild the path contract for each file. The abductive pattern on the mapping, first principles on the pricing simulation, steelmanning on the review, scenario-based on the oral defense -- all of this is already wired in. The pilot works one level above: they validate that the pattern is being followed, intervene when it drifts, adjust at the margin according to the particularities of the file. But they do not reinvent the grammar at every session.

What TITAN does not replace: the pilot's fine-grained knowledge of the agent. The platform orchestrates Opus 4.7, but the pilot must still know Opus 4.7's biases on their files, its recurring blind spots, its stylistic regularities. This knowledge remains empirical, situated, to be built over the long term.

And above all, TITAN does not replace attention during production. Without an attentive pilot, even TITAN converges toward the median. The methodological infrastructure amplifies the pilot -- it does not replace them. This nuance is probably the most important to internalize for the departments investing in agentic platforms in 2026: you do not buy a production line without a pilot, you buy a system that makes the pilot more effective.

Operational consequence

For a department that wants to genuinely extract value from semi-autonomous agentic work in 2026, the diagnosis can be stated plainly. Prompt engineering, post-hoc evaluation, model selection, the deployment of an agentic platform -- all these investments are necessary, but none engages the central competency.

That central competency is steering. And it demands a training of its own, which exists in almost no AI transformation program today. Four elements form its core.

First, the understanding of the agent as both mirror and person -- which presupposes prolonged exposure to the same agent, on real files, with explicit feedback on what that exposure reveals about the pilot's thinking.

Next, the internalization of the asymmetry of obedience -- the hidden cost of each remark, the discipline of focusing on the current phase, the active restraint over the outside comments one is tempted to slip in.

Then, the strategic / tactical / operational grid -- which the pilot must be able to name silently before each intervention, and which they must refuse to mix within a single turn.

Finally, the reading of early-drift signals -- vocabulary that slips, nuance that flattens, epistemic markers that fade -- and the discipline of observing several turns before acting, and of acting at the right level when one acts.

These competencies are not acquired in theoretical training. They are acquired through practice, at length, on real files, with explicit feedback on the errors made. They resemble the training of a mission lead more than that of a tool user. This is probably why, in 2026, they are among the rarest and most differentiating a senior executive can have.

Semi-autonomous agentic work does not eliminate the need for a competent human. It displaces it. Where you once needed juniors who execute and seniors who validate, you now need pilots who think, who observe, who intervene at the right moment and at the right level. The organization that has trained its senior executives in this stance in 2026 will not merely produce better deliverables -- it will have, on complex missions, a structural advantage its competitors will not have seen coming.


TITAN, TenderGraph's cognitive system, applies these best practices — language science, rhetoric, cognitive science — to every bid it works on. And TenderGraph helps teams optimise their pre-sales process along these lines: discover TenderGraph · talk to our team.

Primary sources -- metacognition & sustained attention: Flavell, "Metacognition and Cognitive Monitoring," American Psychologist, 1979. Kahneman, Thinking, Fast and Slow, FSG, 2011. Chabris & Simons, The Invisible Gorilla, Crown, 2010. -- Expertise in dynamic environments: Klein, Sources of Power: How People Make Decisions, MIT Press, 1998. Schon, The Reflective Practitioner, Basic Books, 1983. Endsley, "Toward a Theory of Situation Awareness in Dynamic Systems," Human Factors, 1995. -- Human-machine relationship: Reeves & Nass, The Media Equation, Cambridge University Press, 1996. Bainbridge, "Ironies of Automation," Automatica, 1983 (foundational text on the paradoxical role of the human supervisor). Turkle, Reclaiming Conversation, Penguin Press, 2015. -- Steering and automation: Lee & See, "Trust in Automation: Designing for Appropriate Reliance," Human Factors, 2004. Parasuraman, Sheridan & Wickens, "A Model for Types and Levels of Human Interaction with Automation," IEEE Transactions on Systems, Man, and Cybernetics, 2000. -- Hierarchy of intervention (strategic / tactical / operational): NATO doctrine AJP-01 "Allied Joint Doctrine," Boyd (OODA Loop concept, 1970s-80s). -- Asymmetry of LLM sycophancy: Sharma et al., "Towards Understanding Sycophancy in Language Models," Anthropic, arXiv 2310.13548, 2023. Perez et al., "Discovering Language Model Behaviors with Model-Written Evaluations," Anthropic, 2022. -- Agentic AI practice: Anthropic, "Building effective agents," anthropic.com, 2024. OpenAI, agent practices documentation, 2024-2025.

Tags

#AI#agentic#steering#bid management#cognitive leadership#transformation

Next step

Ready to transform your tender response?

Keep reading

Recommended articles