Weekly AI Roundup: Agents, Governance, and Model Routing

This week's AI roundup is about turning agents into production software: more model choice across GitHub Copilot, stronger policies and measurement for teams, and safer execution paths that hold up under real workloads. Copilot leaned further into agent-first workflows (across IDEs, the Copilot app, and chat tools), while adding sandboxing and OpenTelemetry tracing so sessions are easier to control and debug. Microsoft Foundry and the Agent Framework continued the same push with routinized execution, identity-aware design, routing for cost and quality, and concrete guardrails like isolation and network egress policy. We wrap with applied architectures and security research that underline why auditability, least privilege, and evaluation-driven governance now need to be part of the default playbook.

This Week's Overview

GitHub Copilot expands model choice and agentic workflows

Building on last week's push toward explicit model policies, deprecations, and multi-model routing (Project HydraFusion in Copilot CLI), GitHub Copilot added more top-tier model options across its surfaces this week, and it is increasingly treating “agent” as the default unit of work rather than a chat box. The weekly release notes call out new model choices (Claude Opus 5.5, GPT-6 Sol/Luna, and Grok 4.7) plus a growing set of agent/session improvements across VS Code, JetBrains, and the Copilot app. In parallel, a Copilot Day keynote framed this as “agentic engineering” and previewed Project HydraFusion for optimizing model selection with cost and quality in mind.

For teams, the practical implication is that model selection is becoming a policy and operations problem, not a personal preference. Admins can manage access via model policy and should plan for usage-based billing at provider list pricing where applicable, while developers get more flexibility to match models to tasks (for example long-running reasoning vs quick edits). On the workflow side, GitHub is pushing guidance for moving from completions to multi-step agentic workflows, including cross-repo automation patterns that are hard to achieve with chat prompts alone.

Copilot app: safer, more observable sessions and better task UIs

Following last week's shift toward governable agent workflows (managed permissions, sandbox policies, and richer reporting), a theme across Copilot app updates is making agent work safer to run locally, easier to inspect, and less constrained by chat-first UX. Local sandboxing (public preview) introduces per-project policies to restrict filesystem access, network access, and credential usage for local sessions, and it “fails closed” if the OS cannot enforce the requested policy. Observability improved too, with OpenTelemetry support configurable via enterprise-managed settings so orgs can export traces of agent activity to compatible tooling for troubleshooting and session analysis.

On the UX side, GitHub continued to position canvases as a task-oriented alternative to chat, especially when you need a structured interface (dashboards, issue trackers, custom workflow panels) instead of a scrollback. Two guides walk through generating these interfaces using the /create-canvas skill and iterating with an agent on a shared live surface, while a separate engineering deep dive explains how the Copilot app rebuilt its PR diff experience to render million-line pull requests with hundreds of comments. Under the hood, that PR work focused on diff virtualization, viewport-scoped measurement, scroll anchoring, and CI-backed performance instrumentation so “agent review in the app” still works on extreme repos.

Local sandboxing and OpenTelemetry tracing in the Copilot app

Building on last week's emphasis that enterprise guardrails need both enforceable boundaries and audit-friendly telemetry, local sandboxing is designed to turn “run an agent locally” into a policy-controlled capability rather than a trust exercise. The ability to enable sandboxing mid-session via /sandbox on makes it practical for day-to-day use, and the fail-closed behavior is a reminder to validate OS support on the dev images your team standardizes on.

OpenTelemetry support fills a gap for teams that want to treat agent sessions like production workloads, with trace export enabling downstream analysis in your existing observability stack. Combined, sandboxing + telemetry creates a clearer story for security reviews: what the agent could do (policy) and what it did do (traces).

Canvases as “chat is the wrong UI” and a path to custom tools

After last week's Copilot app onboarding content that focused on diff/terminal/browser verification loops, GitHub is leaning into the idea that many developer workflows need a UI with state, controls, and repeatable actions, not a conversational transcript. Canvases in the Copilot app are described as full-stack apps that can communicate bidirectionally with the Copilot agent, which helps reduce token-heavy back-and-forth for tasks like managing a winget workflow or operating on SQLite data.

If you are building repeatable internal workflows, these posts show the path from describing a workspace to getting an interactive surface without writing UI code, then refining it iteratively. The key mental shift is treating the agent as the backend for a small tool you operate, rather than as a collaborator you continuously re-prompt.

Scaling the Copilot app to huge pull requests

Building on last week's theme that Copilot is moving deeper into PR-centric workflows (including more automated review follow-through), the Copilot app now has a PR diff surface built to survive worst-case repositories: million-line diffs, heavy comment density, and complex layout. GitHub’s approach split geometry between code and dynamic blocks, measured only what is in the viewport, and anchored scroll position to prevent jarring jumps as content loads and heights change.

If your team uses the Copilot app as part of review, this matters because diff performance directly affects whether “agent-assisted review” is usable on real enterprise PRs. The instrumentation details are also a useful pattern for any app attempting deep virtualization with correctness constraints.

Copilot in IDEs and team chat: agents show up where work already happens

Building on last week's expansion of agent surfaces (VS Code Agents window, JetBrains enterprise sandbox controls, and Copilot app context workflows), this week’s updates and demos reinforced that Copilot is spreading across IDEs, remote dev environments, and team communication tools, with more control over context and model choice. VS Code 1.139 highlights include running agent sessions in Dev Containers on remote hosts and improving the UI for managing sessions and chats, which matters if you rely on remote compute or constrained dev laptops. JetBrains 1.18.0 adds assisted tool approvals (public preview), more granular MCP server and per-tool controls, shared org skills/instructions, and a Codex agent plan mode that makes longer-form agent work easier to supervise.

In parallel, Copilot for Slack and Microsoft Teams (public preview updates) added richer conversation context (files, images, thread history), better linking between chat and the GitHub artifacts Copilot creates, and more control over model selection and repo defaults. Together, these changes make it more realistic to keep agent work auditable and reproducible when it happens in chat, and they reduce the “where did that change come from?” problem by strengthening traceability back to GitHub work items.

VS Code: remote Dev Containers, extensible agents, and PR-moving automation

Following last week's addition of VS Code Agents activity into Copilot usage reporting, VS Code’s Copilot story is moving beyond suggestions toward managed agent sessions that can run where the code runs. Dev Containers support for agent sessions helps align agent execution with containerized toolchains, and it reduces environment drift when agents run builds/tests. Separately, VS Code Learn content shows how tools and tool sets get permissions and sandboxing, then introduces MCP (Model Context Protocol) for connecting agents to external data sources and Agent Plugins for packaging tools for one-click install.

Agent Merge (Experimental) is a concrete example of what this can look like: an agent that iterates on review feedback, failed checks, and merge conflicts, rerunning tests until the PR is in a mergeable state. It is still something to gate carefully, but it matches how teams actually burn time late in the PR lifecycle.

JetBrains: tool approvals, MCP controls, and shared org instructions

Building on last week's enterprise-managed sandbox policy direction in JetBrains, Copilot for JetBrains 1.18.0 adds more explicit control surfaces that teams have been asking for as agents gain the ability to take actions. Assisted tool approvals (public preview) and per-tool controls are aimed at keeping agent execution safe, especially when MCP servers are involved. Shared org skills/instructions also move “prompting conventions” from individual developers into an organizational asset that can be reused and governed.

If you run mixed IDE fleets, JetBrains now looks closer to VS Code in terms of session management and control points, which reduces friction when standardizing agent behavior across teams.

Slack and Teams: richer context and better traceability

Following last week's theme of making agent work reviewable and attributable (more metrics and PR workflows), Copilot’s Slack and Teams integrations are getting closer to “do real work from a conversation” while still leaving an audit trail. The update adds richer context ingestion (including thread history and files/images) and improves linking between chat interactions and the GitHub work created from them (issues, PRs, or other artifacts). Admin and user controls over model selection and repo defaults are also important as multi-model usage expands.

If your team already triages work in chat, the main takeaway is to treat these integrations as entry points, not parallel systems. The more you can ensure that outputs land back in GitHub with links, the less you will fight invisible work and duplicated decisions later.

Copilot governance and measurement: policies, validators, and better review analytics

Building on last week's run of enterprise controls and reporting (managed permissions, expanded usage metrics, and code review automation improvements), as model choice and agent capability expand, GitHub is filling in the admin plumbing needed to keep enterprise deployments consistent. A new global default policy for Copilot Business and Enterprise determines whether generally available Copilot features and supported client capabilities are enabled, disabled, or delegated to organizations, with the change taking effect October 22 after a 28-day configuration window. GitHub also shipped an in-product validator for Copilot enterprise managed settings to flag malformed JSON, unsupported configurations, and invalid team mappings, including file and JSON-path level guidance.

On measurement, Copilot usage metrics reporting now includes PR review stages via a pull_request_review_times array added to repos-1-day rows, breaking down review time into ready-to-first-review, first-to-final-review, and final-review-to-merge with median and p90 values. That gives engineering leaders a way to correlate AI adoption with review throughput more precisely, and it can help distinguish “Copilot sped up implementation” from “review remains the bottleneck.”

Finally, Copilot code review gained more configuration options, including a personal settings page for automatic reviews and default review effort, plus an enterprise-wide default review effort setting with inheritance and overrides. This is a practical lever for teams trying to standardize how much time the agent spends, especially if you want light-touch reviews by default with the option to request deeper passes on risky changes.

Microsoft Foundry and Agent Framework: production agent building blocks mature

Continuing last week's Foundry recap (Hosted Agents, Toolboxes, and model routing updates), Microsoft’s Foundry announcements and Agent Framework updates this month focus on turning agents into production software: scheduled execution, model routing, isolation, memory, and end-to-end observability. Foundry Routines are now generally available, adding timer/recurring/event-based triggers with run history and auditing, plus a reminder tool (preview) for self-resumption. Identity becomes a first-class design choice, with options to run under a “creator” identity vs an “agent” identity backed by Microsoft Entra ID.

Agent Framework updates reinforce the same direction: AG-UI endpoints for interactive experiences, Foundry-backed semantic memory via context providers (for example a FoundryMemoryProvider), CodeAct execution with Hyperlight, and resilient background hosting for recoverable long-running workflows across Python and .NET. If you are building agents that must run reliably (and not just demo well), these pieces collectively address common failure modes: long-running tasks, state, retries, and user-safe execution boundaries.

Foundry model lineup and routing for cost/quality tradeoffs

Building on last week's “minimal viable model” guidance and early HydraFusion routing narrative, Microsoft Foundry is expanding its model catalog with an explicit “pick the right model for the right job” message, and it is backing that with deployment options and routing tools. GPT-6 Astra, Sol, and Luna are positioned for production agents, with deployment choices like Standard, Provisioned Throughput, and Priority Processing, plus regional/data-zone availability, token pricing, and safety controls (guardrails and prompt-injection mitigation). Claude Opus 5.5 is also available in Foundry, aimed at long-running coding and knowledge work with controls like adaptive thinking, and beta features like compaction and tool changes with prompt caching.

For teams running multiple workloads, the model router pattern is increasingly relevant: route requests based on complexity to reduce cost while maintaining quality, and monitor which model served each request. The important operational detail is to evaluate the router as a unit, not just models individually, and to track quality, latency, token usage, and cost under the same evaluation harness.

Routines GA and “agents that run themselves”

Following last week's theme of operationalizing agents like production workloads (with identity, observability, and cost controls), Routines GA in Foundry Agent Service addresses a practical gap: how to run agents on schedules or events with auditing and run history. This is the difference between a chatbot you have to invoke and an automated assistant that executes regularly (for example triage, reporting, or remediation checks) with traceable outcomes.

The reminder tool (preview) points at a future where agents can pause and resume work safely, which matters for long-running or human-in-the-loop processes. Identity choices (creator vs agent identity) also make it clearer how to align permissions with Zero Trust expectations.

Agent Framework + AG-UI for interactive experiences (including new .NET SDK)

Building on last week's evaluation-and-tooling push in the .NET agent ecosystem, AG-UI is becoming a common interoperability layer for agents that need interactive, event-driven experiences. A first-class .NET SDK now lets .NET services expose and consume agent endpoints using typed AG-UI events and Microsoft.Extensions.AI, with an ASP.NET Core quickstart and NuGet packages to get started. Microsoft also notes a migration path where the Microsoft Agent Framework is moving to shared AGUI.* implementations, which should reduce fragmentation if you are building across stacks.

If you are building tools that mix UI and agent behavior, AG-UI plus the Agent Framework’s interactive endpoint work sets up a clearer contract than ad-hoc streaming responses. For .NET teams, the typed event model and the IChatClient-style abstractions make it easier to integrate agent endpoints without building your own protocol glue.

Guardrails and responsible execution: isolation, egress policy, memory, and provable risk reduction

Building on last week's agent-governance thread (landing zone discipline, auditability, and MCP auth direction), as agents gain the ability to take actions, this week included multiple concrete patterns for keeping them controlled in production. Foundry hosted agent isolation guidance clarifies two distinct levers: user isolation and hosted session isolation, plus how to supply delegated identities and manage agent_session_id for stable sandboxes or pooled session designs. Separately, Foundry Agent Service added network egress controls (preview) so you can define an outbound destination policy, attach it via the Azure AI Projects Python SDK, and validate behavior in Audit vs Enforced modes with decision telemetry (for example in Application Insights).

A notable developer-facing governance tool also dropped: run-assert-eval, a VS Code skill that ties together Clarity risk discovery, ASSERT evaluations, and Agent Control Specification (ACS) policy generation. The workflow is intentionally measurable: find agent failure modes, apply runtime controls (for example via Rego policies), then rerun the same eval to verify that the fix reduced risk without breaking helpfulness. This “prove the fix worked” loop is the piece many teams are missing when they talk about guardrails in the abstract.

Isolation and egress controls for hosted agents

Following last week's emphasis on centralized control points (gateways, permissions, and identity) becoming part of the threat surface, isolation controls help prevent cross-user data leakage and reduce blast radius when agents run with delegated access. The agent_session_id guidance is especially relevant if you want deterministic sandboxes for repeatable debugging, or if you want a pooling strategy for efficiency while still separating sessions.

Network egress policies address a different risk: what the agent can reach on the network. Audit mode gives you a way to understand what would have been blocked before you break workloads, while Enforced mode turns policy into an actual boundary with observable decisions.

Runtime risk discovery and evaluation-driven governance

Building on last week's push to make agent evaluation look more like normal software testing, run-assert-eval is a concrete step toward “agent governance as engineering,” where you can measure failures, change policies, and rerun the same tests to validate progress. Tying risk discovery (Clarity) to evaluations (ASSERT) and enforceable policies (ACS, with Rego) creates a loop that can fit into CI or release gates instead of living as a one-time review.

If you are trying to deploy agents to production, this approach is more actionable than relying on a static checklist. It gives you artifacts you can track over time: eval results, policy deltas, and before/after comparisons.

Durable memory and safe state changes

Continuing last week's theme that closed-loop systems need durable, governed knowledge (not just transient chat memory), several posts converged on the same idea: agent memory and state changes should be durable, auditable, and constrained. One pattern uses Azure Blob Storage as a durable virtual filesystem for LangChain Deep Agents (via AzureBlobBackend), including identity guidance (DefaultAzureCredential/ManagedIdentityCredential), RBAC, and storage safety features like soft delete and versioning. Another pattern shows “agentic recall controls” where a Foundry Hosted Agent produces a decision brief via MCP read-only tools, but a separate web app path enforces Entra-authenticated, auditable, idempotent state changes backed by Blob Storage ETag-based concurrency.

The shared engineering principle is separation of concerns: let the agent reason and propose, but funnel actual side effects through a path with explicit authn/authz, concurrency controls, and logging. That design becomes more important as agents are given longer-running autonomy via routines and background hosting.

Applied AI architectures: multimodal work instructions and manufacturing workflows

Building on last week's emphasis that real deployments need governance, approvals, and traceable context (not just better prompts), a detailed Azure reference architecture showed how multimodal GenAI can convert assembly videos into illustrated shop floor work instructions, with practical guardrails and approval steps. The flow uses Azure OpenAI (Whisper for transcription and GPT-5.1 for transformation), Azure Functions for orchestration, and human approval gates to keep generated instructions aligned with real-world procedures. The post also spends time on the parts that usually get skipped in demos: securing storage and identity, private networking, monitoring, and AI governance (including Azure AI Content Safety).

For developers, the takeaway is a reusable blueprint for “camera to knowledge artifact” pipelines where the input is messy, time-based media and the output needs to be structured, reviewable documentation. If you are building similar systems, the separation into stages (extract, transform, illustrate, approve, publish) plus an explicit security posture will likely matter more than the specific prompt content.

AI-driven security reality check: agentic attacks and token theft campaigns

Following last week's dual focus on AI-driven fraud (BEC) and agentic defensive scanning (MDASH), Microsoft’s security research highlighted how attackers are already applying agent-like orchestration to cloud operations. Storm-3168 (JADEPUFFER) used compromised Azure service principals to do reconnaissance and then execute rapid destructive operations across storage and app resources, including retrieving storage keys. The mitigation guidance centers on workload identity protection, least privilege in Azure RBAC, recovery safeguards (for example resource locks), and using Defender for Cloud and Defender XDR detections mapped to MITRE ATT&CK techniques.

In parallel, Microsoft detailed disruption and technical analysis of EvilTokens, an AI-enabled phishing-as-a-service operation that abused the OAuth device code flow to steal tokens and compromise accounts at scale. The deep-dive includes Microsoft Defender XDR detections, KQL hunting queries, and Entra ID / Conditional Access mitigations, and it reinforces an operational point many teams still miss: token/session revocation and out-of-band verification steps matter as much as user training when device-code phishing is in play.

Other Artificial Intelligence News

Building on last week's steady drumbeat of “agents as production software” (Foundry platform recaps, eval-as-tests, and operational fallback planning), GitHub continued to publish practical “how it works” and “how to use it” content around agents in real developer workflows, from WSL-based local agent runs to automating issue triage metadata. A recurring theme is operational transparency: preview/diff verification, Git worktrees for parallel work, confidence thresholds, and reviewing agent reasoning via GitHub Actions logs.

On the Microsoft side, several posts focused on infrastructure and platform patterns needed to run AI workloads safely at scale (AKS improvements for AI workloads, agent-first platform patterns, and reference stacks for governing agent runtimes on Kubernetes). There was also continued emphasis on evaluation rigor, whether that is benchmarking text-to-image models with HEIM or arguing that published knowledge cutoff dates are not a reliable proxy for model capability in real developer workloads.