Weekly GitHub Copilot Roundup: Models, Canvases, and Governance

This Week's Overview

Model choices expand (and admins get more knobs)

GitHub Copilot added several new selectable models this week, continuing the shift toward “pick the right model for the task” instead of a single default, and picking up directly from last week's model-policy-and-lifecycle thread (including the MAI-Code-1-Flash deprecation) by expanding the menu that admins now have to govern. OpenAI's GPT-6 Sol and GPT-6 Luna are now available alongside GPT-6 Astra, with GitHub positioning them as different tradeoffs for common coding workloads (and making them selectable in supported Copilot surfaces). Anthropic's Claude Opus 5.5 is now selectable as well, and GitHub notes that the model watermarks text outputs.

Grok 4.7 is rolling out too, rounding out a week where Copilot's model menu grew in multiple directions at once. For teams, the practical takeaway is that model governance matters more as model sprawl grows: access can be controlled via Copilot model policy, and usage-based billing applies at the provider's list pricing when enabled. If you are standardizing across an org, review where model selection appears (IDE vs Copilot app vs chat surfaces) and decide whether developers can switch freely or should stick to approved defaults per repo/org.

Copilot app: canvases, safer local execution, and better observability

The Copilot app updates this week clustered around making agents more usable for real work: a task-oriented UI (canvases), tighter local controls (sandboxing), and better insight into what the agent is doing (OpenTelemetry), building on last week's theme that the app is turning into a sustained workflow surface (diff, terminal, browser) that teams can standardize on. Together, these changes make the Copilot app feel less like “chat that sometimes edits code” and more like a runtime where you can build, constrain, and monitor agent workflows.

Canvases push beyond chat as the default interface

GitHub is leaning into canvases as an alternative to chat when the work is structured (dashboards, trackers, multi-step workflows) or when you need persistent state and controls instead of a scrolling transcript, which follows last week's emphasis on multi-step work in the Copilot app and pulls that workflow out of the transcript and into a purpose-built UI. Canvases are described as full-stack apps that can communicate bidirectionally with the Copilot agent, which is the key difference from “prompt → response”: the UI can both drive agent actions and reflect live results.

For developers, the interesting part is how quickly these can be created and iterated on. The beginners guide shows using the /create-canvas skill to generate a custom workflow UI, then refining it collaboratively on a shared live surface. Separate demos emphasize that you can describe the tool you want (for example a lightweight issue tracker or dashboard) and have Copilot generate the interactive UI without writing UI code, which makes canvases a practical way to wrap repeatable engineering tasks in a purpose-built interface.

Local sandboxing adds “fail closed” controls for agent sessions

Local sandboxing (public preview) adds per-project policies to restrict filesystem access, network access, and credential usage when running local Copilot app sessions, extending last week's move toward enforceable agent guardrails (enterprise managed permissions and JetBrains sandbox policies) from centralized policy into the developer's local runtime. The standout detail is that it “fails closed” if the OS cannot enforce the requested sandbox policy, which matters for teams trying to make local agent execution acceptable under security requirements.

In practice, this gives you a more controlled way to let an agent operate on a repo locally while limiting blast radius, and GitHub even supports toggling it mid-session with /sandbox on. If you have developers experimenting with agents that can run tools, this is the kind of guardrail that can turn “interesting demo” into something you can pilot on production-adjacent code.

OpenTelemetry support turns agent sessions into debuggable systems

The Copilot app now supports OpenTelemetry configuration via enterprise-managed settings, enabling export of agent activity traces to compatible observability tools, which builds on last week's push for better rollout measurement (VS Code Agents usage metrics) by making individual agent runs inspectable, not just countable. That shifts troubleshooting from “paste chat logs” to “follow traces,” which should help when sessions involve tool calls, external systems, or multi-step automation where failures can be intermittent.

If you are rolling Copilot agents out broadly, plan for how you'll use this data: session-level analysis, debugging why an agent made a particular call, and spotting reliability issues (slow tool execution, repeated retries, unexpected call patterns). This pairs naturally with managed policy controls, since tracing configuration can be enforced centrally.

The Copilot app rebuilt PR rendering to handle extreme-scale diffs

GitHub shared a deep dive on rebuilding the Copilot app pull request diff surface to cope with million-line PRs and hundreds of inline comments, a practical follow-on to last week's framing that the app's “diff + validation loop” is the core safety mechanism once agents start generating larger change sets. The approach combines split geometry (separating code from dynamic blocks), viewport-scoped measurement, and scroll anchoring, backed by CI-driven performance instrumentation so regressions get caught early.

If you use the Copilot app to review changes, this work is less about a new feature and more about removing failure modes that show up at scale (slow scrolling, layout thrash, incorrect anchoring when dynamic elements resize). It is a reminder that “agent + diff UI” becomes its own performance-sensitive product once teams rely on it for large refactors or generated changes.

Policy, governance, and org-wide configuration get stricter (in a good way)

Copilot's surface area keeps expanding (more models, more clients, more agents), so GitHub spent part of the week tightening enterprise controls, continuing last week's story that “agent safety” is becoming enforceable through managed permissions and settings rather than best-effort guidance. The theme across updates is fewer “silent misconfigurations” and more predictable defaults, with settings that can be validated and inherited across org structures.

Default enablement policy lands for Business and Enterprise

GitHub announced a new global default policy for Copilot Business and Enterprise that determines whether generally available Copilot features and supported client capabilities are enabled, disabled, or delegated to organizations. This takes effect October 22, with a 28-day window to configure before the new defaults apply.

For admins, this is the kind of change that can surprise teams if they assume “new GA features just show up” (or never show up). Review how your org wants to handle features like Copilot Code Review and supported client capabilities, and decide whether central IT/security sets the baseline or orgs can opt in/out.

Managed settings now have an in-product validator

Enterprise managed settings now include an in-product validator that flags malformed JSON, unsupported configurations, and invalid team mappings, which complements last week's theme of turning governance into day-to-day operations by reducing the “policy looks fine but did not apply” debugging loop. The useful detail is that it provides file- and JSON-path-level guidance, so you can quickly locate what broke and why instead of chasing “policy not applied” symptoms.

If you manage settings as code (or copy/paste policies between environments), this should reduce rollout risk. It also encourages more ambitious configurations (team mappings, fine-grained policies) because the feedback loop is tighter when mistakes are caught before they become confusing runtime behavior.

Copilot Code Review gets more configurable, and metrics get more actionable

Copilot is slowly turning code review into a configurable system rather than a one-size-fits-all feature, building on last week's “agent-assisted PRs” direction (auto-resolution of comments and deeper review analysis) by adding knobs that let teams standardize review behavior instead of leaving it to chance. This week's updates add clearer controls for individuals and enterprises, plus better reporting around the stages of PR review so teams can measure what changed.

Code review configuration expands (personal and enterprise defaults)

Copilot code review updates are now generally available, including a dedicated personal settings page for automatic reviews and a default “review effort” level. Enterprises also get a default review effort setting with inheritance and overrides, making it easier to standardize expectations while still letting teams tune for repo risk.

The practical angle is consistency: if you want Copilot reviews to be present and appropriately thorough, you can set defaults instead of relying on each developer to configure behavior. If you are piloting Copilot review across multiple teams, decide what “effort” means for you (quick checks vs deeper reasoning) and set it at the right scope (enterprise baseline with repo overrides where needed).

Usage metrics now include pull request review stages

The Copilot usage metrics API now adds a pull_request_review_times array to repos-1-day rows, breaking PR review time into ready-to-first-review, first-to-final-review, and final-review-to-merge, which extends last week's expansion of usage metrics beyond basic chat into agent activity by adding a PR-outcomes layer teams can actually optimize against. Each stage includes median and p90 values, which helps distinguish “most PRs are fine” from “a tail of PRs is stuck.”

If you track whether Copilot adoption changes delivery speed, this gives you more precise instrumentation than a single end-to-end PR duration. For example, if “ready-to-first-review” improves but “final-review-to-merge” worsens, you might have a policy or CI bottleneck rather than a review throughput issue.

IDE and CLI agent workflows: Dev Containers, JetBrains controls, and C++ indexing

This week's IDE-focused updates point in one direction: agent sessions are becoming more like repeatable environments (containers, remote hosts, tool approvals) rather than ad-hoc chat, continuing last week's emphasis on governed execution (JetBrains sandbox controls and CLI routing) by making the environment and tool access part of the default workflow. That is good news for teams who need reliability, traceability, and safer tool execution.

VS Code 1.139: agent sessions in Dev Containers (including remote hosts)

VS Code 1.139 includes Copilot updates that make agent sessions easier to run inside Dev Containers, including on remote hosts, building on last week's VS Code Agents and routing discussion by pushing agent runs into the same reproducible environment developers already use for toolchain parity. That matters if your dev environment depends on containerized toolchains or you want the agent to execute in the same environment as CI, not on a developer laptop with a different setup.

The companion deep dive on the inline suggestions model explains why UX and model training are being tuned together in VS Code. GitHub moved inline suggestions to a unified “3-in-1” model combining completions, near-cursor edits, and longer-distance edits, and reports a 10.1% decrease in dismissals after end-to-end optimizations (model + client UX + flighting).

JetBrains 1.18.0: approvals, rewinds, shared skills, and tighter MCP control

GitHub Copilot for JetBrains 1.18.0 adds assisted tool approvals (public preview), message re-editing with rewind, and shared org skills/instructions to standardize how Copilot behaves across a company, extending last week's JetBrains sandbox and diagnostics work into more granular, in-the-moment controls over what tools run and why. It also introduces Codex agent plan mode and more granular controls for MCP servers (Model Context Protocol) and per-tool permissions, plus reliability and UX updates.

For teams using IntelliJ-based IDEs, the approvals and per-tool controls are the headline because they help constrain what an agent can do, which becomes essential once the agent can run tools. Shared org skills/instructions are the other practical piece: they let you encode house style and workflow guidance once rather than repeating it across individual prompts.

Copilot CLI: whole codebase indexing for faster C++ navigation

GitHub Copilot CLI now supports whole codebase indexing (WCI) for C++ code intelligence, using a persistent symbol index via the Microsoft C++ Language Server, which pairs with last week's CLI momentum (multi-model routing and enterprise-default positioning) by making the CLI a more credible “daily driver” surface for large-repo work. The goal is faster “definitions, references, and symbol search” across large repositories, which is where naive on-demand indexing tends to fall over.

If your C++ repo is big enough that “find references” is part of daily friction, persistent indexing can make Copilot-assisted exploration and refactoring more practical. This is a CLI-side capability, so it is worth checking how it fits into your developer environment setup and whether the index persists across sessions and machines in the way you expect.

Copilot across chat platforms: richer context and better traceability

GitHub Copilot for Slack and Microsoft Teams received public preview updates focused on making “chat-to-GitHub work” less lossy, following last week's demo of Copilot Cloud Agent inside Slack/Teams by filling in the details that make those chats usable as real work entry points (context carryover and reliable links back to GitHub artifacts). The integrations now support richer conversation context (including files, images, and thread history), improved linking between chat and the GitHub work created from it, and more control over model selection and repo defaults.

For teams that actually work out of threads, the key improvement is traceability: when a chat request becomes an issue, PR, or other artifact, you want a reliable path back to the source context. Model switching and repo defaults also matter more now that Copilot offers multiple models, since you may want different defaults in chat than in an IDE depending on cost and output expectations.

Developer workflows in practice: WSL agents, automation, and real-world engineering

A few pieces this week focused on how Copilot fits into day-to-day engineering, especially when the “coding” part is just one step in a workflow, and they echo last week's emphasis on validation loops and worktrees as the practical way to keep agent output reviewable. The common thread is moving from one-off completions to repeatable, environment-aware, multi-step routines.

Running Copilot agents inside WSL with worktrees for parallel work

GitHub published a guide showing how to connect the Copilot app to WSL on Windows and run Copilot coding agents directly inside Ubuntu, continuing last week's Copilot app workflow guidance by applying the same worktree-based parallelism and diff/preview verification loop to a Linux-on-Windows setup. The walkthrough highlights using Git worktrees to do parallel feature work, then using in-app preview and diff verification to review what the agent changed before you merge anything.

If your team standardizes on Linux tooling but many developers are on Windows, this is a pragmatic setup: the agent runs where your toolchain runs, and worktrees keep concurrent tasks isolated without juggling branches in a single working directory. The verification loop (preview + diff) is the part to emulate, since it makes agent output reviewable instead of “trust and hope.”

Automating repetitive PR triage and review tasks in the Copilot app

GitHub demoed using the Copilot app to automate a daily PR review workflow: grouping dependency update PRs by risk, checking CI build statuses, and summarizing what is safe to merge, which builds on last week's “agent-assisted PRs” theme by shifting from generating fixes to managing the review queue and its decision signals. This is a good example of where agents pay off, since the value is not “write code” but “collect signals and make a recommendation” across multiple PRs.

If you try this internally, define the inputs you trust (labels, dependency metadata, CI checks) and decide what the agent is allowed to do (summarize vs merge). This kind of automation pairs well with sandboxing and tracing, since it often touches credentials and external systems.

GitHub's own engineering teams use agents for large migrations

GitHub shared how it migrated its Primer-based UI from CSS-in-JS to CSS Modules with feature flags, visual regression testing, and incremental rollouts, then removed sx usage and styled-components across dotcom, echoing last week's point that agents work best inside structured workflows (runbooks, validation, approvals) rather than as standalone generators. The post includes measured SSR and initialization performance improvements and calls out codemods and Copilot coding agents as accelerators for the migration work.

The key pattern here is combining automation layers: codemods handle mechanical changes, while agents help with the messy edges and iterative cleanup. If you are planning a similar migration (CSS stack, framework upgrades, API refactors), this is a useful model for keeping risk down while still moving through a large codebase quickly.

Other GitHub Copilot News

GitHub Copilot Day put a spotlight on “agentic engineering” and how GitHub is thinking about cost/quality tradeoffs at runtime, including an introduction to Project HydraFusion (aimed at optimizing model selection), which connects directly to last week's HydraFusion routing mention in Copilot CLI by framing it as a broader runtime strategy rather than a single-surface experiment. The event framing matters because it connects the week's concrete features (model choice, sandboxing, telemetry, canvases) into a single direction: Copilot as an agent runtime and SDK surface, not just an IDE plugin.

A couple of adjacent reads are worth keeping in mind if you are evaluating agents: Waldek Mastykarz argues that published knowledge cutoff dates are not a reliable proxy for whether a model can work with a given product stack, and recommends workload-based evals and controlled comparisons. On the “Copilot in real debugging” side, Aaron Powell shows using Copilot to help interpret a Visual Studio memory dump investigation, where Parallel Stacks points to thread-pool saturation and Copilot assists in understanding what the dump is telling you.