GitHub Copilot Updates Five Systems at Once

GitHub Copilot updates agent metrics, CodeQL catches prompt injection, and GPT-6 Astra and Gemini 3.6 Flash launch. Learn what this means for business AI.

  • Five releases in one window—and one theme
  • The problem companies keep solving incorrectly
  • What changed: metrics for agent features
  • What changed: protection against prompt injection in CodeQL

Five releases in one window—and one theme

Over two and a half months—from July to September 2026—GitHub rolled out five Copilot updates at once: a new metrics API for agent features, a prompt injection detector in CodeQL 2.26.0, GPT-6 Astra in GA, the GPT-5.6 trio (Sol, Terra, Luna), and Gemini 3.6 Flash.

Individually, they are a routine changelog list. Together, they send a signal: the question “which model should we deploy”

gives way to the question “how do we measure and secure what is already working?”

The problem companies keep solving incorrectly

  1. A typical picture in a company that adopted Copilot six months to a year ago: the tool was distributed to developers, some configured MCP servers, and others created custom agents and slash commands for their workflows—but management has no answer to a simple question: which configurations actually deliver results and which are dead weight sitting unused in the configuration.

  2. Without usage visibility, the licensing budget becomes a matter of taking things on faith.

  3. The same blind spot exists with security: an agent reads external data (a ticket, a PR comment, or file contents), and if it contains an instruction for the model, the agent may execute it against the user's intent.

  4. Until recently, this could only be detected manually.

What changed: metrics for agent features

  1. GitHub expanded the usage metrics API with totals_by_skill, totals_by_custom_agent, and similar fields for MCP servers, slash commands, and plugins.

  2. The fields are available in enterprise and organization reports: daily per-user and aggregated reports, 28-day per-user reports, and aggregated 28-day reports in day_totals.

  3. For the first time, organizations can see which custom agents and MCP servers are actually being used and which have long been untouched.

  4. For a CTO, this metric is the basis for cutting unnecessary integrations and investing in those that deliver TTU (time to use): the faster a specific configuration gives a developer a working result instead of merely occupying a place in the list of available tools, the more valuable it is.

Assess where AI can deliver impact in your process

What changed: protection against prompt injection in CodeQL

CodeQL 2.26.0 added the js/system-prompt-injection query, which detects cases where untrusted external data reaches an LLM system prompt. Support for Kotlin 2.4.0 arrived as well. For teams already embedding LLMs into their products rather than merely using Copilot as an assistant, this is the first static analyzer of its kind in mainstream tooling—previously, injection checks were performed manually or not at all.

A seemingly simple outcome: “we found a vulnerability before production”

- relies on serious engineering work: data flows between external input and the system prompt must be modeled precisely to avoid drowning in false positives.

What changed: multiple models instead of one

Instead of one “best” model, Copilot now offers a choice for each task. GPT-6 Astra in GA is designed for long agent sessions: planning, validating intermediate steps, and self-checking before declaring a task complete—fewer cases where an agent reports “done” with barely functioning code. The GPT-5.6 trio—Sol for complex reasoning across large codebases, Terra as the everyday workhorse, and Luna for light, low-cost tasks—acknowledges the obvious: not every refactoring job requires a top-tier model.

Overpaying in tokens for a flagship call on a trivial task drains the budget. According to GitHub, Gemini 3.6 Flash showed a higher completion rate and better token efficiency than 3.5 Flash in agent workflows—a concrete metric, not vague claims about being “smarter.”

One story behind five releases

Measurement (metrics API), protection (CodeQL), and choosing the right tool for each task (the model lineup) are the building blocks of a mature AI environment for development.

A one-off deployment of a trendy assistant falls short of this kind of environment.

Without metrics, a company cannot see where it is losing money on unused integrations.

Without injection analysis, it cannot see where it is losing control of the agent.

Without a deliberate choice of model for each task, a company overpays for capacity where it is unnecessary and cuts costs where savings come at the expense of quality. In projects where we deploy an LLM & Security Gateway and an MCP environment for a client at KT.Team, these three layers—usage metrics, injection checks, and routing requests to a model with the right cost and capability—become part of the infrastructure from day one, rather than being added after an incident or budget audit.

The business does not measure the fact that Copilot with GPT-6 is installed somewhere.

The business measures the outcome: a specific backlog feature reached production faster, without data leaking where it should not.

Conclusion

The model race will continue—new versions will follow GPT-6 Astra and Gemini 3.6 Flash in the coming months. Connecting a model first is not an advantage. Visibility is: can the company see how the model is used and stop it when it goes wrong? Those who build measurement and protection can benefit from each new model release within weeks. Those who do not keep paying for potential that will never become an outcome.

Discuss the article: GitHub Copilot updated five systems at once.…

Enter your email or phone number so we can get back to you.

Send via: