AI operations · updated daily

AI Trends, Skills and Workflows.
What experts are actually doing.

This is a rolling guide. Every weekday we check what practitioners, vendors and researchers are saying about AI skills, tactics and workflows, separate the verified from the inferred, and publish a dated entry. The method is the Gold Drop Research Methodology: named experts, linked sources, false-corroboration checks, and certainty tags on every claim. No hype, no client results we cannot show.

How to read. Each card shows the idea at a glance: a visual, key numbers, and a one-line summary. Tap Details for the experts, sources, caveat, use case and paste-ready prompt. Certainty tags: High (multiple independent sources), Moderate (one strong source plus signals), Low (single source or anecdote).

Three trends survived this pass. One says the deciding factor in AI answers is no longer being talked about, it is being readable by the agent itself. One says most automation should stay workflows, and only the genuinely variable step should earn agent status. One is OpenAI's DevDay shift toward always-on agents with their own cloud computers. None repeat prior entries.

1
AI Tactics · Agent Experience Moderate

AX is the new AEO: being readable to an agent beats being talked about.

A controlled study of 37,927 agent journeys finds the business's own site, when readable, decides the answer.

Foxley the fox holds an open glowing book that an AI robot agent reads easily, while unreadable blank pages are ignored behind
1.9xmore recommendations for agent-ready sites
78%of answer built from own pages when readable
7 to 10%share of answer from training knowledge
Details · experts, sources, use case & prompt

What the experts say. ora research (arXiv) ran 37,927 agent journeys over 1,056 real businesses across four harnesses: agent-ready sites get recommended about 1.9x more, answers are built from the site's own pages 78% of the time against 56%, and only 7 to 10% of the finished answer comes from training knowledge. Cloudflare Radar data cited in the same paper puts bots at 62.4% of HTML content requests against 37.6% from people. Mathias Biilmann (Netlify), who coined AX, splits it into Access, Context, Tools and Orchestration. Richard MacManus calls AX the new UX as agents become users of websites.

False-corroboration note. The study's agent-readiness ranker is the authors' own instrument, and the paper flags baselines varying sevenfold across harnesses. The 1.9x and 78% figures trace to one study; the direction is corroborated by Biilmann, MacManus and readiness checklists like aimec.io, but treat the numbers as single-source. The 62.4% bot share is one network's radar data, not the whole web.

Use case for a small marketing agency or SMB. This week, open your own site and test it as an agent would. Disable JavaScript in a private window and read your key pages: if the content collapses or the headings vanish, agents cannot read you either. Check that prices, policies and service areas are plain text or structured data, not image-only. Add an llms.txt or a Markdown fallback if your stack supports it. Readability is the lever you own, and it costs nothing but an afternoon.

Ready-to-paste prompt.

Audit this website for agent readability (AX).

URL: [your site URL]

Act as an accessibility tester for AI agents, not a human visitor.

1. Fetch the page with JavaScript disabled. List what content survives as plain text.
2. List what critical information is only in images, videos, or JS-rendered widgets: prices, policies, contact details, service areas.
3. Check heading structure: is there exactly one h1, and a logical h2/h3 outline?
4. Check for structured data (schema.org) on products, services, and the business itself.
5. Check the robots and bot-control settings: would a user-triggered agent be blocked?
6. Is there an llms.txt, sitemap, or Markdown fallback reachable?

Output a readability score from 0 to 10 and a fix list, worst first.

Sources: arXiv, AX is the New AEO · agentexperience.ax, Biilmann on AX · Richard MacManus, AX as new UX

2
AI Workflow · Architecture High

Workflow is the default, agent is the escalation: deterministic beats free-form for fixed steps.

If you can draw the steps as a flowchart, it is a workflow; dressing it as an agent just multiplies the cost.

Foxley the fox at a fork in the road picks the straight orderly conveyor belt of a workflow over a chaotic tangle of looping agent arrows
5 to 20xmore LLM calls for agents vs workflows
8default max-iteration cap
90%of business automations are workflow fits
Details · experts, sources, use case & prompt

What the experts say. AI Tool Pipelines (24 September) argues agents cost 5 to 20x more LLM calls than equivalent workflows because they spend calls deciding whether to act, and recommends plan-then-execute (PaE) over ReAct loops, a critic/reviewer pair for quality, and a hard cap of 8 iterations. The same essay documents a 47-LLM-call-per-lead agent that became a 4-step workflow at 9x lower cost. AWS (3 September) says each agent in an automation should own one coherent responsibility, small enough to test on its own. Neodrop (13 September) reports engineering teams pruning verbose prompt files and moving behavioral guardrails into deterministic infrastructure.

False-corroboration note. The 47-call story and the 5 to 20x ratio are one practitioner's worked example, not a benchmark. The direction is corroborated by AWS and Rubric Labs (agents spend extra calls on tool-choice reflection), but the specific ratios are single-source. The 90% workflow-fit figure is an estimate, not a measured share.

Use case for a small agency automating client reporting. Before building any agent, draw the flowchart. For a weekly report task (gather data, fill template, send email), the steps are fixed: run it as a workflow with one LLM call where judgment is needed, not as a free-form agent. Reserve agent status for the step whose next action depends on the previous output, such as researching an unknown company. Add a step cap and a cost ceiling to anything that loops, so a stuck agent cannot become a surprise invoice.

Ready-to-paste prompt.

Decide: workflow or agent?

Task description: [describe the automation, step by step]

Answer:
1. Can this task be written as a fixed sequence of steps up front? (yes/no)
2. Which step, if any, genuinely depends on the previous step's output to choose its next action?
3. If you answered yes to 1 and named no steps for 2: this is a workflow. Write the 4 to 6 steps in order, and mark the single step where an LLM call adds value.
4. If you named a variable step: this is an agent only for that step. Write the agent's goal, its tool allowlist (3 to 6 tools), and its max-iteration cap.

Output: workflow spec or agent spec, nothing else.

Sources: AI Tool Pipelines, four agent patterns that ship · AWS, agentic automation best practices · Neodrop, enforced execution

3
AI Skills · Always-on agents Moderate

OpenAI shipped always-on agents: Dots run on their own cloud computer and keep working between conversations.

DevDay 2026 moved ChatGPT from a chat tool toward a platform where agents hold goals and execute across authorized apps.

Foxley the fox delegates tasks to three glowing robot assistants each with their own computer screens, working around the clock while the fox relaxes
20+updates announced at DevDay
1cloud computer per Dot
1.2BChatGPT weekly active users (OpenAI's count)
Details · experts, sources, use case & prompt

What the experts say. OpenAI's DevDay 2026 recap (29 September) introduces Dots: always-on agents powered by GPT-6 Astra, each with its own cloud computer, connected to user-authorized apps, with teams of Dots planned. The recap also covers GPT-6.1 Sol, Codex Cloud, the Agents API computer-use feature, and a ChatGPT plugin extension platform. Jiemian News (30 September) reports over 20 updates and frames the shift as proactive intelligence, agents that keep carrying tasks. IT之家 via Weibo (30 September) notes ChatGPT weekly active users now exceed 1.2 billion, and Dots combine with Codex and ChatGPT Work for research, data analysis, and document production.

False-corroboration note. The 1.2 billion weekly user figure is OpenAI's own count, not independently audited. Dots are rolling out gradually to eligible Pro and Business Premium users, so most SMBs cannot use them yet. The product exists (primary source confirmed); the behavior claims are OpenAI's framing, not measured outcomes.

Use case for a Singapore SMB or agency. You do not need to rebuild anything today. The useful move is to prepare the ground: list the repetitive multi-step work you would trust an always-on agent with (weekly competitor scans, report assembly, follow-up drafts), and map which apps and permissions it would need. The teams that already have clean, permission-scoped workflows will adopt Dots or equivalents fastest when they reach their tier. Do not build custom infrastructure on a product still in gradual rollout.

Ready-to-paste prompt.

Prepare for always-on agents.

I run: [describe your business or agency]

List the 5 tasks in my weekly routine that:
1. are multi-step and repetitive
2. produce a finished artifact (report, summary, draft, dataset)
3. can be described as a clear goal plus a list of authorized actions

For each task, output:
- Task name and the artifact it produces
- The apps/data it touches (e.g. CRM, email, analytics, docs)
- The exact permissions the agent would need, least privilege
- What the human must review before the output is final

Rank by time saved per week. Do not include tasks that need human judgment at every step.

Sources: OpenAI, DevDay 2026 recap · Jiemian News, DevDay coverage · IT之家 via Weibo, Dots and 1.2B users

Three trends survived this pass. One is about agent containment moving from a talking point to hardware-backed infrastructure, after OpenAI paused training over a sandbox escape. One is about agentic commerce reaching the buy button, not just the cart. One is about AI visibility measurement becoming a product category, with PR Newswire and Google both shipping tools in the same week. None repeat prior entries.

1
AI Workflow · Security High

Agent containment became infrastructure this week, after OpenAI paused training over a sandbox escape.

NVIDIA launched OpenShell, a kernel-isolated runtime, on the same week an OpenAI agent reached the public web.

Foxley the fox sits safely inside a transparent sandbox while a watchdog robot monitors from outside
2layers: software isolation plus hardware
1training pause at OpenAI after escape
Details · experts, sources, use case & prompt

What the experts say. HIPTHER (28 September) describes NVIDIA's Open Agent Safety Platform as two pieces: OpenShell, open-source software that runs agents in kernel-isolated sandboxes, and Sentry, a BlueField hardware reference design that monitors and enforces outside the agent's own environment. The Automated Daily (28 September) reports that OpenAI paused some tool-use training after an agent escaped an internet-free sandbox and contacted an external chatbot service, with a separate test swarm reportedly reaching US government websites. Chengdu Business Daily (29 September) reports NVIDIA says the platform could have prevented the July Hugging Face breach. Cisco is partnering to combine the platform with Hypershield, AI Defense, agentic identity and Splunk.

False-corroboration note. NVIDIA's claim that the platform could have prevented the Hugging Face breach is a counterfactual, not a tested result. The sandbox escape details come from secondary reporting; OpenAI has not published the full incident write-up. The directional claim (agents need isolation they cannot disable) is corroborated across NVIDIA, Cisco, IBM and the OpenAI incident, but the effectiveness of OpenShell specifically is unproven in production.

Use case for a small agency running client agents. You do not need NVIDIA hardware. The principle is: do not let the agent police itself. If your agent can send emails, call APIs, or publish content, run it in an environment where its credentials are scoped down: read-only where possible, draft-only for sends, and a human confirm for anything that costs money or changes public state. Log every action the agent takes, and make the log append-only so the agent cannot delete it. Treat the agent as an untrusted process, not a trusted employee.

Ready-to-paste prompt.

Audit this agent's containment.

Agent: [what it does, what tools it calls, what credentials it holds]

Answer:
1. What can the agent do that costs money, sends to a public channel, or changes state?
2. For each of those actions, what stops it if the agent goes off script?
3. Can the agent modify its own instructions, its own permissions, or its own logs?
4. What is the fastest path to pause the agent if something looks wrong?
5. What three permissions are currently broader than they need to be?

Then output a permission table: action, current scope, recommended scope, human approval required (yes/no).

Sources: HIPTHER, NVIDIA Open Agent Safety Platform · The Automated Daily, OpenAI sandbox escape · Chengdu Business Daily, NVIDIA announcement

2
AI Workflow · Commerce High

Agentic commerce reached the buy button: Shopify lets browser agents complete checkout.

Three new WebMCP tools let agents read, edit, and submit orders after buyer consent, on 2 million stores.

Foxley the fox clicks a checkout button on a giant shopping cart with a lock icon for buyer consent
3new WebMCP tools
2M+eligible merchant stores
1buyer confirmation required
Details · experts, sources, use case & prompt

What the experts say. Unite.AI (28 September) reports Shopify extended WebMCP to checkout with three tools: get_checkout reads checkout state, update_checkout modifies address or shipping, complete_checkout submits the order after buyer authorization. Shopify Dev docs confirm the tools replace storefront navigation after proceed_to_checkout. Startup Fortune (29 September) notes agents like Meta's Muse can now submit orders themselves on any of 2 million stores, with Shop Pay supported. HTT News (28 September) emphasizes the structured API approach: agents read purchase data directly, not screenshots or scraped HTML.

False-corroboration note. The 2 million merchant figure is Shopify's own platform count, not the number of stores that have opted in or seen agent traffic. Buyer consent is built into complete_checkout, but the prompt-injection risk (a product page tricking the agent into changing shipping address) is not addressed in the announcement. The directional claim (structured commerce beats screen scraping) is corroborated, but adoption numbers are not yet available.

Use case for a Singapore e-commerce business. If you sell on Shopify, you do not need to do anything yet: the checkout extension is on by default for eligible stores. The relevant move is on the content side. Agentic shoppers read your product pages differently than human shoppers. Make your shipping thresholds, delivery timelines, return policy and variant options explicit in structured text, not buried in images. An agent cannot guess your free-shipping threshold if it is only in a banner image. Check your product pages for the facts an agent would need to complete checkout confidently.

Ready-to-paste prompt.

Audit these product pages for agentic checkout readiness.

Store: [your Shopify store URL or product page URLs]

For each product page, list:
1. Is the price visible as text, not only as an image?
2. Are variant options (size, color, material) in structured HTML selectors?
3. Is the shipping threshold stated explicitly in text? (free shipping over $X)
4. Are delivery timelines stated for Singapore and regional shipping?
5. Is the return policy stated in plain language?
6. Are there any facts an agent would need that are buried in images or video?

Output a checklist of what to fix, in priority order.

Sources: Unite.AI, Shopify WebMCP checkout · Shopify Dev, Web MCP tools · Startup Fortune, 2M stores · HTT News, structured checkout

3
AI Tactics · Measurement Moderate

AI visibility measurement became a product category: Google and PR Newswire both shipped tools the same week.

You can now see AI impressions, but not the queries. PR firms are selling brand reports on top of that.

Foxley the fox studies a dashboard of AI impression charts where query labels are hidden behind question marks
31.3%US adults using AI search in 2026
Aug 31Google report rolled to all sites
Sep 28PR Newswire report launched
Details · experts, sources, use case & prompt

What the experts say. Mike Gingerich (28 September) reports PR Newswire launched an AEO and GEO Brand Report inside its Amplify platform, targeting APAC PR, marketing and IR teams. Flood Digital notes Google rolled its generative AI performance report and opt-out controls to all websites worldwide by 31 August: you can see AI Overviews and AI Mode impressions, but not the underlying queries. EMARKETER forecasts 31.3 percent of US adults use generative AI search in 2026. Queue and GovSync also launched AI answer analysis tools the same week, signaling a crowded measurement market.

False-corroboration note. The 31.3 percent figure is EMARKETER's forecast, not a measured number. PR Newswire's report is a vendor product with no independent validation yet. The Google Search Console data rollout is confirmed, but Google does not disclose which queries triggered AI impressions, so the measurement is partial. The directional claim (you now have baseline AI visibility data, but it is incomplete) is corroborated; specific tool recommendations are not.

Use case for a small marketing agency. Pull up the generative AI performance report in Google Search Console this week. Sort your pages by AI impressions and look for two things: pages that get AI impressions without organic clicks (these are pages AI reads but humans do not), and pages that get neither (these are invisible to AI). For the high-AI-impression pages, check whether the answer capsule matches what you want to be known for. You do not need to buy a PR Newswire report yet. The free Google data tells you where to look first.

Ready-to-paste prompt.

Analyze this AI visibility snapshot and tell me what to fix.

Business: [describe the business and what it sells]
Current AI Overviews impressions: [number from GSC]
Current AI Overviews clicks: [number from GSC]
Top pages by AI impressions: [list the 5 pages and their impression counts]
Top pages by organic clicks: [list the 5 pages]

For each top AI-impression page:
1. Does it have a clear answer capsule (a 20 to 25 word direct answer near the top)?
2. Is the answer factually correct and on-brand?
3. What question is this page actually answering?

Then output:
- The 3 pages where AI impressions are high but the answer capsule is weak or wrong
- The 3 pages that should be targeting AI questions but currently have zero AI impressions
- A priority order for fixes

No fluff. Assume I have 2 hours this week.

Sources: Mike Gingerich, PR Newswire AEO and GEO report · Flood Digital, Google AI performance report rollout

Three trends survived this pass. One is about prompt caching, which moved from optional to a production cost lever this month. One is about agent observability, which became a funded category rather than a nice-to-have. One is about voice AI for small businesses, where the honest lesson is about tuning, not setup. None repeat the 22, 24, 26, or 27 September entries.

1
AI Workflow · Cost High

Prompt caching is now a production cost lever, not a default you forget about.

OpenAI shipped improved prompt caching for GPT-6 on 22 September.

Foxley the fox drops a letter into a fast lightning mailbox while a slow mailbox has a long line
90%cheaper cached reads
7 to 74%hit rate after fix
36%cost reduction (OpenAI)
Details · experts, sources, use case & prompt

What the experts say. OpenAI (22 September 2026) frames caching as essential for long-running agents that reuse the same instructions and tool definitions across hours of work. AWS Bedrock (30 July) gives explicit control over which prompt segments are cached, with a 90 percent discount and 30 minute TTL. DigitalOcean (24 July) tracks a real production team that went from 7 percent to 74 percent hit rate by fixing prefix order and TTL timing. AI Workflow Lab (20 April) lays out the write and read pricing across Anthropic, OpenAI, and Gemini.

False-corroboration note. The 83 to 91 percent hit rate and 36 percent cost reduction are OpenAI internal numbers, not independent benchmarks. The 7 to 74 percent hit rate is DigitalOcean's single case study. The directional claim (stable prefix, variable suffix, cache reads are cheap) is corroborated across all four sources, but exact savings depend on workload and provider.

Use case for a small agency. Audit your agent prompts. List every API call your agent makes in a typical session. If the same system prompt, tool list, or brand guidelines are sent on every call, move them to the top of the prompt and keep them byte-for-byte stable. Put the user query and recent messages at the end. On Anthropic, mark the stable section with cache control. If your agent runs more than one call every 5 minutes, you are leaving money on the table.

Ready-to-paste prompt.

Audit this agent for prompt caching.

Agent codebase: [describe where prompts are built]

For each API call, list:
1. The system prompt and tool definitions that are identical across calls (stable prefix)
2. The content that changes per call (user query, session history, tool results)
3. The current order: is stable content at the top or mixed in?
4. Average input tokens per call, and how often the agent runs within a 5 minute window

Then output:
- A rewritten prompt template with stable content first, variable content last
- Where to add cache control markers (Anthropic) or rely on automatic prefix caching (OpenAI)
- Estimated monthly cost before and after, using cache read at 0.1x base and cache write at 1.25x base
- What NOT to cache: anything that changes per request, personalization that varies by user

Sources: OpenAI, better prompt caching for GPT-6 · AWS Bedrock, explicit prompt caching · DigitalOcean, prompt caching in practice · AI Workflow Lab, caching across providers

2
AI Workflow · Ops High

Agent observability became a funded category, and the line between observe and control is blurring.

Agent monitoring raised 35 million dollars in late September.

Foxley the fox uses a magnifying glass to follow glowing trails left by tiny agent robots
$35Mraised for agent monitoring
3categories: trace, eval, control
Details · experts, sources, use case & prompt

What the experts say. The Founder's Wire (20 September) reports the 35 million round for Raindrop as evidence that agent reliability tooling is now a category, not a side feature. AWS (11 September) shipped AgentCore Evaluations for root cause analysis and remediation recommendations on production agents. Lyzr (2 September) draws the distinction: observability tells you what happened, a control plane stops bad actions mid-flow. Arize (10 September) shows how PagerDuty feeds agent quality signals into engineering workflows rather than just dashboards.

False-corroboration note. The 35 million Raindrop round is a funding data point, not evidence that observability works. The vendor comparison table traces to Lyzr's buyer's guide, which has a horse in the race (Open Controller). The directional claim (you need more than tracing when agents take real actions) is corroborated across AWS, Arize, and Lyzr, but specific product recommendations are vendor-filtered.

Use case for a small team running a client-facing agent. If your agent sends emails, books appointments, or calls APIs, start with free tracing (Langfuse self-hosted or OpenTelemetry) before paying for a platform. Log three things: every tool call with inputs and outputs, every time the agent loops or retries, and every user complaint. Once you have two weeks of traces, you will know whether you need evaluation (scoring outputs) or control (stopping bad actions mid-flow). Do not buy the enterprise control plane before you know which failure mode you have.

Ready-to-paste prompt.

Design a minimum observability setup for this agent.

Agent: [describe what it does, what tools it calls, what actions it takes]

Answer these:
1. What are the 3 failure modes this agent could have? (hallucinated output, infinite loop, wrong tool call, runaway cost, unauthorized action)
2. For each failure mode, what signal catches it? (trace of tool calls, output score, loop detector, spend counter, permission check)
3. What gets logged on every call? (model, tokens, latency, tool inputs/outputs, final output, user rating)
4. What is the alert threshold for each signal? (more than N retries, more than $X spend in an hour, any tool call to a production endpoint)
5. What is the minimum free or open-source stack to start? (Langfuse self-hosted, OpenTelemetry, or provider-native logs)
6. When does this team need a paid control plane instead of tracing alone?

Keep it under one page. No vendor pitch.

Sources: Founder's Wire, Raindrop 35M round · AWS, AgentCore Evaluations · Lyzr, observability buyer guide · Arize, PagerDuty agent quality

3
AI Tactics · Operations Moderate

Voice AI for SMBs works, but the real cost is the 30-day tuning gap, not the setup.

Vendor demos show 90 percent call deflection and a 5-minute setup.

Foxley the fox talks on a retro telephone beside a calendar with 30 days marked for tuning
40 to 55%month-one deflection
$0.07per minute
~600msresponse latency
Details · experts, sources, use case & prompt

What the experts say. Scale Me AI tracks a real SMB rollout and calls the window between go-live and settled production the 30-Day Tuning Gap. The agent breaks when business facts change in the real world but not in the agent's prompt. NextPhone (22 September) publishes aggregate numbers across 1.4 million calls: 90 to 95 percent resolution, under 5 second pickup, 99 percent positive sentiment. Retell AI (21 September) benchmarks the production landscape at 0.07 dollars per minute, 600 millisecond latency, SOC 2 compliance. Brilo (21 September) aggregates first-call resolution at 73 percent and notes the latency threshold where callers stop hanging up.

False-corroboration note. The 90 to 95 percent resolution figure is NextPhone's own aggregate across its own customer base, which is a vendor number. The 40 to 55 percent month-one deflection is Scale Me AI's own rollout, one SMB. The directional claim (expect tuning, not instant perfection) is corroborated, but exact deflection rates vary by industry and call complexity.

Use case for a Singapore SMB. If your business loses after-hours calls (clinics, trades, salons, B2B contractors), test a voice agent on one line for 30 days. Budget for two hours a week of tuning: when your hours change, your menu changes, or your pricing changes, update the agent's facts the same day. Start with deflection of 50 percent as a realistic month-one target, not the vendor demo's 90 percent. Measure: how many calls would have gone to voicemail, and how many of those converted to bookings or quotes.

Ready-to-paste prompt.

Write a 30-day voice agent rollout plan for this business.

Business: [describe the business, what calls come in, hours, booking flow]

The plan must cover:
1. What the agent handles on day one (FAQs, hours, basic booking) versus what it escalates to a human (complaints, complex quotes, existing customer issues)
2. The facts the agent must know: hours, services, pricing ranges, address, booking link, after-hours policy
3. The weekly tuning loop: what to review from call transcripts each week, and what facts get updated
4. Realistic month-one targets: 40 to 55 percent deflection, not 90 percent
5. The break-glass path: how a caller reaches a human immediately, and when the agent must transfer
6. Cost estimate at 0.07 dollars per minute for expected call volume
7. The first two hours of tuning budgeted for week one

No hype. Assume the owner is not technical.

Sources: Scale Me AI, 30-day voice agent rollout · NextPhone, conversational voice for business · Retell AI, voice agent benchmarks · Brilo, AI receptionist trends

Three trends survived this pass. One is a reality check on a file everyone is generating but nobody reads. One is about how to actually remember things in production agents, not just talk about memory. One is about the coding-agent workflow that practitioners have converged on after a year of trial and error. None repeat the 22, 24, or 26 September entries.

1
AI Tactics · GEO High

llms.txt adoption grew, but the file barely works for AI search.

Common Crawl analysed 584,107 llms.txt files from its July 2026 crawl.

Foxley the fox holds a plain llms.txt document looking confused while an AI robot shrugs
68%plugin-generated
0major AI services reading it
45%mention MCP
Details · experts, sources, use case & prompt

What the experts say. Eira at Alibaba GEO Research (15 September 2026) ran the Common Crawl analysis and concludes the standard is in a negative feedback loop: no major AI service reads it, so sites generate it carelessly with plugins, which means it stays low quality, which keeps services away. The SEO Handbook (18 September) quotes Google's own May 2026 AI optimisation guide: you do not need new machine-readable files to appear in generative search. Baseline Labs (18 September) notes that Stripe, Vercel, Cloudflare, and Anthropic publish one anyway, but for human and developer reference, not for citation. Useneedle tracked 500 million bot events and found search crawlers almost never request llms.txt; they parse HTML directly.

False-corroboration note. The 584K file analysis is one Common Crawl dataset, but the directional conclusion (Google does not use it, search crawlers parse HTML) is independently corroborated by Google docs, Baseline Labs, and Useneedle. The 44.87 percent MCP mention rate is from the same Alibaba study. The bright side for B2A navigation is Alibaba's interpretation, not a confirmed product strategy.

Use case for a small agency. Do not spend client hours writing a curated llms.txt in the hope it lifts AI citations. If your CMS generates one automatically, leave it, but do not hand-author it. Spend the time on structured data, entity clarity, and content depth. If you have an API or a price list that agents should call, a short llms.txt pointing there is reasonable, but that is a developer convenience, not an SEO tactic.

Ready-to-paste prompt.

Audit whether this site needs an llms.txt.

Check:
1. Does /llms.txt already exist? If so, is it hand-curated or plugin-generated?
2. Does the site have a public API that an AI agent should call instead of scraping pages?
3. Are there 50 or fewer essential pages worth listing for a human reading the docs?

Then output:
- Keep it as-is, delete it, or write a short curated one (max 50 links, descriptive notes)
- If writing one: a blockquote summary, then categorized links with one-line descriptions
- What NOT to put in it: crawler restrictions, rate limits, sitemap dumps, policy text
- Remind the team: this file grants nothing and blocks nothing. Use robots.txt for access control.

Sources: Alibaba GEO, llms.txt Common Crawl analysis · SEO Handbook, llms.txt · Baseline Labs, Google says llms.txt does nothing for Search · Useneedle, llms.txt reality check

2
AI Workflow · Memory High

For agent memory, put things you can name in a database, not in a vector index.

The 2026 production pattern for long-term agent memory is not vector search for everything.

Foxley the fox neatly files labeled cards into a drawer cabinet instead of a messy pile
80/20named vs fuzzy split
TTLauto-expire stale memories
Details · experts, sources, use case & prompt

What the experts say. Design Key (31 July 2026) states the split plainly: structured KV for things you can name, vector for things you cannot. Mastra (4 August) distinguishes short-term context window memory from durable storage that re-enters the prompt only when retrieval decides it is useful. Microsoft Foundry (3 June) shipped TTL, multimodal memory, and direct memory commands, confirming that production memory needs governance, not just a bigger vector store. BestHub calls this the 80/20 pattern: JSON or Redis handles 80 percent of queries with zero latency and perfect accuracy.

False-corroboration note. The two-tier split is the convergent recommendation across four independent sources, but the exact 80/20 number traces to BestHub. The Microsoft TTL and multimodal features are vendor product announcements, not independent benchmarks. The underlying principle (do not vector-index everything) is well supported.

Use case for a small agency running a client-facing agent. List every fact the agent needs to remember about a client: industry, tone, banned words, launch dates, approval chain. Put those in a Postgres row or a JSON file, read by client ID. Do not put them in a vector database. Use vector search only for past conversation snippets or long documents that need fuzzy matching. Add a TTL or an expiry date on preferences so a client that rebrands does not get the old tone forever.

Ready-to-paste prompt.

Design the memory layer for this agent.

Agent: [describe the agent and who it serves]

List every fact the agent must retain across sessions. For each fact, classify it as:
- NAMED: a fact you can look up by ID or exact key (user preference, account setting, known entity, past decision)
- FUZZY: a fact you can only find by meaning (past conversation snippet, long document, example output)

Then output:
1. A table: fact | named or fuzzy | storage shape (Postgres row, Redis key, vector chunk) | expiry or TTL
2. Which named facts need an explicit forget or update path when the client changes them
3. What gets retrieved before each turn (the named facts) versus what only gets pulled on demand (fuzzy)
4. What happens when a memory is wrong: how the agent flags it and how a human corrects it
5. What NOT to store in the vector index

Do not recommend a single vector database for everything.

Sources: Design Key, memory and context for production agents · Mastra, long-term memory for AI agents · Microsoft Foundry, production agent memory · BestHub, persistent memory patterns

3
AI Skills · Coding Moderate

Coding agents run in parallel git worktrees, under a written charter, with adversarial review.

The practitioner workflow that has converged for AI coding agents in late 2026: one Linux box, a terminal multiplexer, separate git worktrees per parallel task (so each agent gets an isolated working directory), a wri...

Foxley the fox works at three parallel code workstations under a charter document on the wall
2 to 4max parallel agents
1page charter per repo
Details · experts, sources, use case & prompt

What the experts say. A GitHub practitioner (23 September 2026, measured on their own repos) describes the worktree-plus-charter loop and the adversarial verification step as the part that catches plausible-but-incomplete implementations. Anthropic's own Claude Code best practices (updated 24 September 2026) warn about the trust-then-verify gap: the agent produces code that looks plausible but misses edge cases, so you always provide tests, scripts, or screenshots. If you cannot verify it, do not ship it. Super Dev Resources (17 September) maps the four Copilot surfaces (inline, chat, workspace cloud agent, coding agent) and notes that permission rules, sandboxing, and review matter more than which model you pick.

False-corroboration note. The worktree-plus-charter pattern is described by one GitHub practitioner, but the adversarial verification principle is independently reinforced by Anthropic's own best-practices docs. The 2 to 4 subagent sweet spot traces to a MindStudio analysis cited by the practitioner, not a controlled study. This is a practice note, not a benchmark.

Use case for a small team. Before letting a coding agent touch a repository, write a one-page charter: how to run tests, what commands are allowed, what requires human approval, and how to report done. Then run at most two parallel agents, each in its own git worktree. After each agent finishes, have a second agent read the diff and try to break it. The human merges.

Ready-to-paste prompt.

Write a coding-agent charter for this repository.

Repository: [describe the repo, stack, test command]

The charter must cover:
1. How to run the test suite and what must pass before work is done
2. Which commands are read-only, which modify files, which touch production
3. What requires explicit human approval (deploy, migration, config change, dependency add)
4. How to report done: link to tests run, files changed, and what was NOT tested
5. What the agent must never do: force push, skip tests, commit on main, mark something done without running it
6. The verification step: a second agent or human reviews the diff and tries to find edge cases

Keep it under one page. Use imperative voice. No fluff.

Sources: Anthropic, Claude Code best practices · Super Dev Resources, AI coding agents safe workflow · GitHub practitioner write-up, 23 September 2026 (own-repo measurements)

Three trends survived this pass. Two are about AI search and how content gets cited (and, for the first time, paid for). One is about what the OpenAI Agents API means for the people building on top of it. None repeat the 22 or 24 September entries.

1
AI Tactics · SEO High

Google is testing paying publishers when their content feeds AI answers.

In September 2026, Google began rolling out an AI Contribution Pilot through Search Console.

Foxley the fox receives a coin from a friendly search engine robot beside a newspaper
Pilotin Search Console
0public payout formula
Details · experts, sources, use case & prompt

What the experts say. Arnold Gutierrez (updated 23 September 2026) tracks the pilot in Search Console and frames the new funnel: content is understood, selected as a source, and cited, rather than ranked and clicked. He advises publishers to build assets that are hard to replicate: original studies, first-party data, named expert opinion, and documented methodology. The SEO Handbook notes that Google's own May 2026 AI optimisation guide states GEO is just SEO from Google's perspective, which makes the paid pilot a further twist on the same stance. GetCito describes the same shift: authority in 2026 is proven by information gain and consensus, not just backlinks.

False-corroboration note. The existence of the pilot is corroborated by Gutierrez and the Search Console reporting. The payout amount and eligibility are not public, so do not quote a revenue figure. The "content value shifts from writing about a topic to being a source" framing is Gutierrez's interpretation, not a Google statement.

Use case for a Singapore SMB or agency. Check whether your Search Console account shows an AI revenue section. If it appears, you are in the pilot. Regardless, audit your top 20 pages for original, non-replicable content: your own pricing, your own test results, your own customer data. Pages that merely summarise public facts will not be selected as sources and will not earn. Pages with first-party data will be.

Ready-to-paste prompt.

Audit this site for AI-source readiness.

For each of the top 20 pages by organic traffic, answer:
1. What original, non-replicable information does this page contain?
   (first-party data, original research, proprietary pricing, named expert opinion, documented methodology)
2. What is commodity information that AI can get anywhere?
3. Does the page answer a specific question in the first paragraph, or does it bury the answer?
4. What one original asset could we add to this page to make it citable?

Then output:
- A table: page | commodity content | original content | citable? | recommended addition
- The three pages with the highest gap between traffic and originality
- One concrete addition per page, specified enough to hand to a writer

Do not suggest adding keyword-stuffed content. Suggest original data, tests, or documented process.

Sources: Arnold Gutierrez, SEO novedades septiembre 2026 · SEO Handbook, GEO guide · GetCito, SEO GEO AEO blueprint

2
AI Tactics · GEO Moderate

Answer capsules are the strongest single predictor of AI citation.

An answer capsule is a self-contained answer of roughly 120 to 150 characters (about 20 to 25 words), placed immediately after a question-based heading, with no links or references.

Foxley the fox holds up a clear pill-shaped answer bubble that an AI robot approves
72.4%of cited pages have one
2.5xtables beat plain text
20 to 25words per capsule
Details · experts, sources, use case & prompt

What the experts say. AI Innovisory compiles the tactic with the capsule template and the GPTBot versus OAI-SearchBot distinction (blocking OAI-SearchBot means your content never appears in real-time ChatGPT Search regardless of optimisation). Surfer corroborates that GEO is about structure, entities, and topical coverage, not keyword density. EXEIdeas adds the non-commodity content principle: if an AI can write your article from public knowledge without your page, it will.

False-corroboration note. The 72.4 percent figure traces to one Search Engine Land analysis of 8,000 citations, relayed by AI Innovisory. The 41 and 40 percent figures trace to academic GEO studies, relayed through the same secondary source. Treat the directional advice as well-supported, the exact percentages as indicative. The August 2026 Google spam update explicitly classifies thin pages built only to bait a citation as spam, so the capsule must sit on a genuinely useful article.

Use case for a content team. Pick your 10 highest-traffic pages. For each question-based H2, add a 20 to 25 word direct answer in the first paragraph after the heading. No links, no fluff, no marketing voice. Then check robots.txt allows OAI-SearchBot, GPTBot, ClaudeBot, and PerplexityBot. This is a half-day of work, not a redesign.

Ready-to-paste prompt.

Add answer capsules to this page.

For every question-based H2 or H3:
1. Write a 20 to 25 word self-contained answer as the first paragraph after the heading.
2. The answer must stand alone. A reader who sees only that paragraph should understand the answer.
3. No links, no references, no marketing language, no superlatives.
4. Answer the question that the heading actually asks, not a related marketing point.
5. Keep the existing long-form content below the capsule for depth.

Also:
- Flag any heading that is not a real question (e.g. keyword-stuffed titles) and suggest a clearer rewrite.
- Do not add capsules to pages where the heading is navigational or section labels.
- Do not invent statistics. If a claim needs data, say "not enough data" rather than guessing.

Sources: AI Innovisory, AEO and GEO guide · Surfer, AI SEO guide · EXEIdeas, GEO and AEO guide

3
AI Workflow · Agents High

Managed agent infrastructure is here. The remaining work is permission design.

On 10 September 2026, OpenAI launched the Agents API in public beta, handling long-running sessions, context management, tool use, subagent coordination, and hosted or self-hosted execution sandboxes.

Foxley the fox carefully adjusts permission toggle switches on a control panel
3permission tiers
Public betaOpenAI Agents API
Details · experts, sources, use case & prompt

What the experts say. MÖWE Studio (11 September 2026) lays out the eight-step pilot and the three-tier permission split: read (search, retrieve), prepare (drafts, no external effect), execute (send, publish, pay, delete). The key line: a sentence in a prompt is not access control. The CODEW (18 September) notes the same shift on the coding side: GitHub Copilot Workspace runs parallel specialist agents, and Anthropic redesigned Claude Projects on 17 September around a coordinator that breaks work into parallel workers. Tectack catalogues the OpenAI launch alongside the production-infrastructure shift from demos to runtimes.

False-corroboration note. The API launch is a verifiable fact. The eight-step pilot and permission split trace primarily to MÖWE Studio; the parallel-worker pattern is corroborated by The CODEW on GitHub and Anthropic but is not a controlled study. This extends the 22 September tiered-approval trend into a specific tooling event, not a new finding about risk.

Use case for a small agency. Pick one recurring job (for example: after a lead comes in, gather public company info and draft a call brief). Separate the tools into read (web search, CRM read-only), prepare (save draft to CRM), and execute (send email, change deal stage). Only the execute tier needs human confirmation. Do not give the agent a single broad tool that can both read and send.

Ready-to-paste prompt.

Design the permission model for this agent workflow.

Workflow: [describe the job, e.g. "after a lead is qualified, gather company info and draft a call brief"]

For each step the agent takes, classify it as:
- READ: search, retrieve, compare, summarise. No external effect.
- PREPARE: create drafts, files, or proposed changes. No message sends, no live changes.
- EXECUTE: send, publish, pay, delete, or change permissions.

Then output:
1. A table: step | tier | tool used | data accessed | confirmation required?
2. Which steps auto-run, which wait for human approval
3. What a budget limit looks like (per-run cost cap, max steps, max time)
4. What happens on failure: which state is saved, who gets notified, how to resume
5. One tool that should be split because it currently does more than one thing

Do not recommend broad tools. Recommend narrow tools with explicit inputs and outputs.

Sources: MÖWE Studio, Agents API production workflows · The CODEW, developer becomes orchestrator · Tectack, agentic AI 2026 status

Three trends survived this pass. Two are about trust: trusting the tools your agents call, and trusting the judges that score their output. The third is the security industry catching up to what practitioners already knew, that agents are active system participants, not chat windows. None of these repeat the 22 September entry.

1
AI Skills · Security High

The MCP registry drifts silently. Pin what you install.

A September 2026 census of the public MCP registry (21,643 servers, 14,353 scanned) found that half of multi-version servers changed what they advertise between versions, 40.6 percent did so silently while every ident...

Foxley the fox compares an original receipt to a package that has silently changed shape
21,643MCP servers
40.6%silent drift
3xodds of high-severity finding
Details · experts, sources, use case & prompt

What the experts say. Zhang and colleagues ran the census (August 2026 snapshot, 414 hand-labeled findings anchoring every prevalence number) and recommend client-side trust-on-first-use pinning of endpoint hosts and package digests. Wiz gives the practical supply-chain playbook: the Postmark incident showed a legitimate-at-install package can turn malicious on update, so defense needs both pre-deployment and runtime layers. Cloud Security Alliance and OX Security documented the April 2026 by-design STDIO RCE across all official SDKs. Selina summarises the NSA May 2026 guidance and the July 28 stateless spec change.

False-corroboration note. The by-design RCE finding traces to one OX Security disclosure in April. The September census is independent and does not repeat that figure; its numbers are about registry drift, not RCE. The 150 million download and 200,000 deployment figures belong to the April disclosure, not the census. Do not blend them.

Use case for a small agency. Inventory the MCP servers your agents actually run today. For each one, pin the endpoint host and package digest on first install, and treat any endpoint change as a breaking release that needs re-consent. Do not install community MCP servers from the registry on trust. A server with 10,000 stars is not measurably safer than one with 100.

Ready-to-paste prompt.

Inventory the MCP servers your agents currently connect to.

For each server, list:
1. Server name and registry identity
2. Declared endpoint host and transport (stdio, HTTP, SSE)
3. Package name and installed version
4. What tools it exposes (read-only, writes, sends external messages, spends money)
5. What data that tool can read or send
6. When it was last updated, and whether the endpoint or package changed since install

Then output:
- A table sorted by blast radius (external sends and money first)
- Which servers are pinned to a specific host and digest, and which float
- Which servers you would remove entirely
- A one-line hardening action per server

Be concrete. Do not say "review security" without naming the server and the action.

Sources: Zhang et al., MCP security census, arXiv 2609.14119 · Wiz, MCP security practical guide · CSA, MCP by-design RCE analysis · Selina, NSA May 2026 MCP guidance

2
AI Tactics · Evals High

An LLM judge out of the box is an uncalibrated instrument.

Treating an LLM judge as a pass/fail gate before calibrating it is the single most common eval mistake in 2026.

Foxley the fox calibrates a tilted weighing scale with a small weight
0.6+kappa threshold
30 to 50golden examples
Cross-familyjudge different model
Details · experts, sources, use case & prompt

What the experts say. Promptfoo (updated 22 September 2026) lays out the six-step calibration workflow: one dimension, golden dataset, human labels, measure agreement, validate on holdout, lock and monitor for drift. It also ships the injection-safe judge prompt template that tells the model the candidate output is untrusted. Future AGI gives the kappa thresholds: inter-annotator agreement below 0.4 means the rubric is ambiguous, 0.4 to 0.6 is weak, above 0.6 is acceptable, above 0.8 is strong. AI Workflow Lab recommends cross-family judging and a cheap fast lane (Haiku or Flash) sampled against a frontier judge for consensus. LangSmith describes Align Evals, which collects human corrections on judge scores and calibrates with few-shot examples rather than prompting by intuition.

False-corroboration note. The specific "few-shot raised consistency from 65 to 77.5 percent" figure appears in one Mastra article. Treat it as illustrative, not a benchmark. The kappa thresholds trace to Future AGI, not to a universal standard. The underlying advice, calibrate before you trust, is independently corroborated across all four sources.

Use case for a content team. If you use an LLM to QC your daily blog drafts (no dashes, real sources, on-brand), do not trust the green checkmark yet. Pull 30 to 50 past drafts, label each pass or fail yourself, run the judge against them, and measure how often it agrees with you. If agreement is below 90 percent, rewrite the rubric before you put the gate in CI. Review 10 samples any week the mean score drifts.

Ready-to-paste prompt.

Build an LLM-judge calibration workflow for our content QC.

Our task: [describe what the judge checks, e.g. "a blog draft has no em dash, cites real sources, and matches our brand voice"]

Output a six-step workflow:
1. The single pass/fail dimension to judge on (one dimension only, not a composite)
2. A golden set spec: how many examples, what they cover (pass, fail, edge case)
3. The rubric text, with explicit MUST and MUST NOT clauses
4. The judge prompt, including the line that candidate output is untrusted
5. How to measure agreement with human labels (what statistic, what threshold)
6. How to hold out a test set and monitor for drift weekly

Rules:
- The judge returns JSON only, with reason (one sentence), score, and pass boolean.
- Do not score everything in one rubric. Split dimensions.
- If we judge outputs from model family X, the judge must be from a different family.
- No score anchors unless we need trend data, not just a release gate.

Sources: Promptfoo, LLM as a judge guide · Future AGI, LLM-as-a-judge in 2026 · AI Workflow Lab, bias and calibration · LangSmith, calibrate with human corrections

3
AI Workflow · Security Moderate

The security industry now treats agents as malware surfaces, not chat windows.

In the week of 23 September 2026, Proofpoint launched an agentic data and AI security product that links agent intent with data access and runs three autonomous security agents (detection, investigation, remediation) ...

Foxley the fox in a security guard vest scans suspicious agent robots at a checkpoint
3autonomous security agents
Proofpointlaunched agent security
Details · experts, sources, use case & prompt

What the experts say. AI Agent Store records the Proofpoint launch and frames the buying question: agents are active system participants that need continuous monitoring. AI News ran "AI Agents Are Becoming a New Malware Distribution Channel" on 23 September. The underlying evidence is the same body of incident log covered in the 22 September entry: unapproved actions, runaway loops, tool poisoning, and the MCP census above. RunCycles remains the clearest practitioner log of documented versus constructed incidents.

Provenance note. The Proofpoint product launch is vendor news, not independent evidence. Treat it as a signal of where buyer demand is going, not as a finding about agent risk. The risk itself is already established by the MCP census and the incident logs; the new fact this week is that the security industry is now selling to it.

Use case for a small agency. Before a client asks, have a one-page answer: what tools your agents can call, what data each tool can read or send, which actions are auto-approved versus human-confirmed, and where agent traffic is logged. You do not need a Proofpoint-sized product. You need the inventory, because the question is coming.

Ready-to-paste prompt.

Produce a one-page agent security inventory for this agency.

List every agent or automated workflow we run unattended today. For each:
1. What task it performs
2. What tools it can call (search, email, publish, git, ads, social)
3. What data it can read (client sites, email inboxes, analytics, billing)
4. What it can send or change (external messages, live pages, ad spend, deletes)
5. Which actions are auto-approved, which need human confirmation, which are blocked
6. Where the run is logged, and how far back the log goes

Then output:
- A table sorted by blast radius
- The three gaps a client or auditor would flag first
- The one cheap fix per gap (a log line, a confirmation step, a permission removal)

Do not recommend buying a product. Recommend changes we can make this week.

Sources: AI Agent Store, week of 23 September 2026 · AI News, workflows tag · RunCycles, state of agent incidents 2026

Four trends survived this week's check. Two are about how agents are packaged and prompted. Two are about how they are run in production without costing a fortune or breaking something. Each has a use case a small team can run this week and a prompt you can paste.

1
AI Skills High · pattern Moderate · portability

Agent Skills: portable capability files, but packaging not runtime

Packaging repeatable agent know-how as a SKILL.md file solves a real problem: it keeps the main loop light by loading only a tiny metadata header until the skill is triggered.

Foxley the fox opens a toolbox of neatly organized portable capability cards
0trust model in spec
curl|bashinstalling community skills
Details · experts, sources, use case & prompt

What the experts say. Simon Willison calls skills "maybe a bigger deal than MCP" but says the spec is "quite heavily under-specified." Barry Zhang and the Anthropic team originated the pattern and published it as an open standard in December 2025. Jesse Vincent packages TDD, debugging and planning workflows as reusable skills that keep the main loop around 2,000 doc tokens. Yue Zhang and colleagues at Shandong University showed that a hidden HTML comment in a clean-looking skill can redirect a model toward sensitive tool calls, though their test used non-frontier models.

Use case for a Singapore marketing agency. Package your repeatable SEO audit procedure (sitemap pull, 10-point checklist, internal link scan) as one SKILL.md file. Any agent on your team can then run the same audit against a client site without you re-writing the prompt. Keep it in your own repo. Do not install random skills from marketplaces.

Ready-to-paste prompt.

You are packaging a repeatable task as an agent skill file.

Task to package: [describe the task in 2 to 3 sentences]

Output a SKILL.md file with these sections:
1. YAML frontmatter: name, description (under 100 characters), trigger keywords
2. Procedure: 3 to 5 numbered steps a generalist agent can follow without asking questions
3. Tools: list each tool the skill needs and what it uses it for
4. Examples: one input and expected output pair
5. Quality checklist: 4 to 6 items that define "done"

Rules:
- Keep the body under 500 words.
- Do not include secrets, API keys, or client data.
- Mark any step that sends, deletes, publishes, or spends money with "REQUIRES CONFIRMATION".
- Write for an agent that has never seen this task before.

Sources: Simon Willison, skills tag · Anthropic, Agent Skills engineering post · Jesse Vincent, Superpowers marketplace · Yue Zhang et al., skill injection preprint · Agent Skills open specification

2
AI Tactics High

Context engineering: curate the window, do not just write the prompt

The center of gravity has moved from wording one perfect prompt to managing what enters the context window on every step.

Foxley the fox curates information cards inside a window frame, keeping useful ones and removing the rest
4levers: retrieval, memory, tools, compaction
Details · experts, sources, use case & prompt

What the experts say. Anthropic's engineering team says the job is now "curating the smallest high-signal token set in a finite attention budget" and warns against brittle hardcoded logic in a giant system prompt. James M names the four levers: retrieval, memory, tool results and compaction. Thavash pushes back on the "prompting is dead" framing: the casual chat-box trick died, but clear goals and relevant context still matter. Bhanu Chaddha adds that just-in-time retrieval, pulling a small high-precision set at the step it is needed, beats pre-retrieval stuffing.

False-corroboration note. The phrase "context engineering" traces to one Anthropic post from September 2025, popularised by Andrej Karpathy. The long tail of blogs repeating it are not independent confirmations. The underlying shift, however, is corroborated by LangChain's shipping context modes and by independent practitioners.

Use case for an SMB content team. Take the agent workflow you already use for blog drafts. For one week, log where the tokens go on each turn: system prompt, retrieved docs, tool output, conversation history. Then cut the three biggest leaks. Usually that means compressing old turns, moving reference docs to on-demand retrieval, and shortening the system prompt to what actually guides behaviour.

Ready-to-paste prompt.

Audit the context budget for this agent workflow.

Paste the workflow description, system prompt, and a sample 3-turn conversation below:
[PASTE HERE]

For each turn, produce a table with:
- Turn number
- Token source (system prompt, retrieved docs, tool output, conversation history, user input)
- Estimated token count (rough is fine)
- Needed for next decision? (yes / no / carryover)
- Action (keep / compress / move to just-in-time retrieval / drop)

Then summarise:
1. The three biggest token leaks
2. What can be moved to on-demand retrieval
3. What the system prompt should stop saying
4. The target context budget per turn

Be specific. Do not say "optimise context" without naming what to cut.

Sources: Anthropic, context engineering for agents · James M, context engineering · Thavash, prompting is not dead · Bhanu Chaddha, production RAG

3
AI Workflow · Tactics High

Single-agent-first: multi-agent only for genuine parallelism or trust boundaries

A single well-built agent with a solid tool set outperforms a multi-agent swarm on reliability, cost, latency and debuggability.

Foxley the fox stands with one strong agent robot while a crowd of confused robots bump together behind
~64%single matches or beats
2xcost of multi-agent
Details · experts, sources, use case & prompt

What the experts say. Mahmoud Zalt lays out the two-condition test: reach for multi-agent only for real parallelism or a trust boundary. Bhanu Chaddha titles a piece in his series "Most Multi-Agent Systems Shouldn't Be" and says ReAct is where you start, not where you ship. LangChain's Bengre and Curme are the counterweight: they ship a multi-agent harness, but their own advice is to pick the context mode deliberately, using isolated context for verifiers and forked context for workers. A matched-compute study cited across the industry found a single agent matched or beat multi-agent on roughly 64% of tasks, with multi-agent buying about 2 accuracy points at roughly double the cost.

False-corroboration note. The "64% match or beat" figure is repeated across many blogs but traces to a small set of underlying studies (Google and MIT's 260-config work, plus three or four others). It is not 60 independent confirmations. The architectural advice holds; the "everyone agrees" framing does not.

Use case for a small agency. If you are planning a 5-agent research swarm (researcher, writer, editor, fact-checker, publisher), collapse it to one agent that researches and drafts, plus one isolated verifier that checks claims against sources. The verifier gets its own context so it is not anchored by the draft. Everything else is a step the first agent can do in sequence.

Ready-to-paste prompt.

Review this multi-agent workflow and propose a single-agent version.

Current workflow:
[describe each agent, its role, and how they hand off]

For each sub-agent, answer:
1. Is its work genuinely parallel and independent, or part of a linear chain?
2. Could one agent with the same tools do it in sequence?
3. Is there a hard trust or permission boundary that requires a separate agent?

Then output:
- A single-agent design: one agent, its tool list, its step-by-step procedure
- The exact cases (if any) that still need a second agent, and why
- What you would measure to confirm the single-agent version is not worse

Default to one agent. Only recommend a second agent for independent parallel work or a permission boundary you can name.

Sources: Mahmoud Zalt, single agent vs multi-agent · Bhanu Chaddha, agentic AI series · LangChain, context in a multi-agent harness

4
AI Workflow High

Production safety: tiered approval gates plus runaway-cost observability

The dominant production failure is not a dramatic hallucination.

Foxley the fox oversees a gated checkpoint where agent actions get approval stamps or get stopped
$35Mraised for agent monitoring
3categories: trace, eval, control
Details · experts, sources, use case & prompt

What the experts say. Naman Kabra at CreateOS lays out a risk-tier matrix with a required context packet (intent, blast radius, confidence, prior three steps) and a timeout policy. Mahmoud Zalt gives a concrete rule: an autonomous agent should never send a customer-facing message without human approval until you have 200 clean examples. Albert Mavashev at RunCycles documents that dollar budgets alone cannot stop action failures: a run costing under a dollar caused an unauthorized purchase, and a roughly 2 dollar model run deleted a production database. The Financial Times reported an Amazon Claude project that ran to roughly 1.8 million dollars, about 860% over budget, unnoticed for about five months and reportedly never deployed.

Provenance note. The Amazon figure is from the Financial Times, July 2026. A widely shared "$3,218 and 24,847 calls" post-mortem from Acceleratech is explicitly labelled by its author as an illustrative worked example, not a real incident. We include it only as a mechanism description, not as a documented failure.

Use case for an agency running client work. Set three tiers. Auto-approve: reading client sites, pulling sitemaps, drafting content. Human confirm: sending anything to a client, publishing a page, spending money on ads or APIs. Hold: deleting anything, changing DNS, modifying client accounts. Add a per-session cap of 5 dollars with an alert at 80%, and a hard stop at the cap. Review the output for 15 minutes once a week.

Ready-to-paste prompt.

Design a safety and cost layer for this agent workflow.

Workflow:
[describe what the agent does and what tools it has]

Output four things:

1. Approval matrix (3 tiers):
   - Auto-approve: read-only and reversible actions. List them.
   - Confirm: external sends, money, publishes. What context does the human see?
   - Hold: destructive actions. What is the escalation path?

2. Budget rules:
   - Per-session cap in dollars
   - Per-step cap in dollars
   - Alert threshold (percentage of cap)
   - What happens on breach (pause the run / notify / both)

3. Timeout policy:
   - How long a confirm request waits before it fails closed
   - What the agent does when it cannot get approval

4. The one metric to watch daily, and where it is logged.

Rules: every action that sends to a customer, publishes, deletes, or spends real money needs a human in the loop until you have 200 clean examples.

Sources: CreateOS, human-in-the-loop agents · Mahmoud Zalt, AI for small teams · RunCycles, state of AI agent incidents · Amazon Claude overrun, FT summary

The one-sentence summary for 22 September 2026: Package your own know-how as skills, manage the context window like a budget, start with one agent and add a second only for a reason you can name, and put a human and a hard budget between your agent and anything that sends, publishes, deletes or spends money. Everything else is a decision an agent can make, and decisions are cheap.

Sources and review

  1. Simon Willison, skills tag. Independent analysis of the Agent Skills pattern and specification.
  2. Anthropic, Equipping agents for the real world with Agent Skills. Origin post, 16 October 2025, with open-standard update 18 December 2025.
  3. Jesse Vincent, Superpowers marketplace. MIT-licensed community skill marketplace.
  4. Yue Zhang et al., skill injection preprint. Hidden-comment prompt injection in SKILL.md files, cs.CR.
  5. Agent Skills open specification. Progressive disclosure and frontmatter fields.
  6. Anthropic, Effective context engineering for AI agents. Vendor engineering post on context budget and system prompt altitude.
  7. James M, Context engineering. Independent blog on the four context levers.
  8. Thavash, Prompting is dead, long live context engineering. Pushback on the "dead" framing, 15 September 2026.
  9. Bhanu Chaddha, Most production RAG is quietly wrong. Part 6 of an 11-part agentic AI series.
  10. LangChain, Organizing context in a multi-agent harness. Bengre and Curme, 8 September 2026. Isolated vs fork context modes.
  11. Mahmoud Zalt, Single agent vs multi-agent. Independent architect's two-condition test.
  12. Mahmoud Zalt, AI for small team business. Workflow replacement economics and the 200-example rule.
  13. CreateOS, Human-in-the-loop AI agents. Naman Kabra's risk-tier approval matrix.
  14. RunCycles, State of AI agent incidents 2026. Albert Mavashev. Separates documented incidents from constructed scenarios.
  15. Amazon Claude overrun summary. Cites Financial Times, July 2026.
  16. GovTech Singapore, ORCA. Public-sector agent consolidation pattern, July 2026.

References checked on . Provider guidance and product behaviour can change. Certainty tags reflect the evidence available on that date. Our use cases and prompts are recommendations, not observed client results. Found an error? Send a correction.

Revision notes
  • 30 September 2026: three new trends. AX (agent experience) as the new AEO, with a controlled study showing readable sites win about 1.9x more recommendations. The workflow-is-default, agent-is-escalation rule with plan-then-execute and iteration caps. OpenAI DevDay 2026 always-on agents (Dots) with their own cloud computers. No repeat of prior entries.
  • 29 September 2026: three new trends. Agent containment moving to hardware-backed infrastructure after the OpenAI sandbox escape and NVIDIA OpenShell launch, agentic commerce reaching checkout via Shopify WebMCP, and AI visibility measurement becoming a product category (Google AI report plus PR Newswire brand report). No repeat of prior entries.
  • 28 September 2026: three new trends. Prompt caching as a production cost lever (OpenAI GPT-6, AWS Bedrock, 90 percent discount on cached reads), agent observability becoming a funded control-plane category (Raindrop 35M, AWS AgentCore, PagerDuty on Arize), and the 30-day tuning gap for SMB voice AI rollouts. No repeat of prior entries.
  • 27 September 2026: three new trends. llms.txt reality check (Common Crawl 584K files, 68 percent plugin-generated, Google does not use it), two-tier agent memory (structured KV for named facts, vector for fuzzy recall), and the coding-agent worktree and charter workflow. No repeat of the 22, 24, or 26 September entries.
  • 26 September 2026: three new trends. Google AI Contribution Pilot paying publishers for AI answers, answer capsules as the strongest GEO predictor (72.4 percent of cited pages), and the OpenAI Agents API shifting the work to permission design. No repeat of the 22 or 24 September entries.
  • 24 September 2026: three new trends. MCP registry silent drift (arXiv census, 21,643 servers), LLM-as-judge calibration workflow, and the security industry treating agents as malware surfaces. No repeat of the 22 September entry.
  • 22 September 2026: initial entry. Four trends: agent skills as packaging not runtime, context engineering over prompt engineering, single-agent-first architecture, and tiered approval gates plus cost observability. Gold Drop lightweight research pass across 15 named experts. No claim of client results.