Skip to main content

Capabilities Agentic Engineering

Security Gates for AI-Written Code

Coding agents produce code faster than teams review it. They also install packages, edit settings, and run commands before a pull request exists. Security controls have to run where that happens.


Provectus adapts your existing security workflow for AI-assisted development. We look for what the current tools miss and make the checks part of the merge process, so they cannot be quietly skipped.


The shift

There is more code now, and less of it gets read.

Before anyone opens a pull request, a coding agent may have written the feature, installed a package, and run commands on a developer's machine. Reviewers see the result, but not always everything that led to it. Hooks, plugins, and permission files are part of that path, and they are easy to miss in a conventional code review.

We add checks at the points where agents read files, change code, introduce dependencies, and prepare work for merge. That gives reviewers a record of what ran and what it found.

  • While the agent works.

    Check permissions and sensitive inputs when an agent reads a file, changes code, or adds a dependency.

  • Before merge.

    Run the required checks on the change and return findings the author can act on.

  • When the toolchain changes.

    Review updates to agent tools and settings, then verify that the controls still work.

What we need

Read access is enough to start.

Nothing is installed and the pipeline stays unchanged until you have reviewed the baseline and approved the next stage.

  • Access

    Read access to the repositories, full git history, CI configuration, and the coding agent's configuration files.

  • Existing controls

    Audits, scanner configs, ignore lists, accepted-risk notes, whatever exists. A missing one counts as a finding.

  • Owners

    One engineer who knows the pipeline. An owner for each credential that turns up.

  • Authority

    An administrator who can apply a repository ruleset. Someone named to sign off accepted risks.

The path

Enforce checks based on the evidence collected.

The baseline informs the remediation plan. Once the checks have been tested on a repository, they can be made mandatory and rolled out further.

01 · Measure

Establish the baseline.

We scan the repository history, review the controls already in place, and measure how many scanner alerts are false positives.

What runs

  • We run the same secret-scanning rules across the full history of every repository. We report suspected credentials and do not test whether they still work.
  • The full-history scan covers every repository. The deeper audit, across twelve security dimensions, is scoped to the repositories you name as most important, usually one or two in the first phase, and we verify unexpected results by hand. One of those repositories is where the gates are built and tested before they are ported to the rest.
  • We also record which default branches require passing checks, where those checks can be bypassed, and what the agent configuration allows to run on a developer's machine.

Deliverables: Baseline record

  • Leak count per repository, still present and history only
  • A rotation list with an owner per credential
  • Audit scores, verified by hand
  • The enforcement map for the whole organization
  • A locked scorecard for the next phase

Who says no

No one. The stage is read-only. If the scan comes back clean, we say so.

02 · Triage

Fix the immediate problem and prevent a repeat.

The remediation brief assigns an owner and records both the fix and the control needed to keep the issue from returning.

What runs

  • Rotate first, clean second. Deleting a credential from code does not cancel it; every clone still has it. Owners rotate, then the code is cleaned. History is never rewritten.
  • The fronts. Findings are grouped by the kind of work they require: secrets, AI-written code, dependencies, the AI toolchain, or process and governance. Each one includes an immediate fix and a longer-term control.
  • The accepted-risk register. Deliberate decisions, written down with their reasoning, so audits stop re-flagging them.
  • The enforcement ask. A ruleset requiring the security checks, bypass list empty, drafted for the administrator before any gate is built. The price is fixed, for one repository.

Deliverables: Remediation brief

  • Rotation list with owner verdicts
  • Five-front map, fix today and stay fixed per finding
  • Accepted-risk register, ruleset draft, fixed price

Who says no

Rotation happens in the issuing system, by the credential's owner. Nothing is gated until the brief is signed.

03 · Gate

Test the gates on one repository.

We implement the checks in a real workflow, tune the false positives, and verify how each check behaves when it finds a problem or receives bad input.

What runs

  • The gates. Secrets, agent, code, dependencies, toolchain, posture. Each one sits where the change moves: the commit, the agent's tool call, the merge, the manifest, the pull request.
  • The ignore list. The usual one excuses whole files, and that hides a real token beside a known false alarm. Every entry names its false alarm and excuses one value.
  • The code review. It runs on the agent's own diff before commit and is installed as a required step. Its catch rate is not yet measured, and we say so.
  • The proof. Before a gate is accepted, we run three tests on it. It has to catch a finding we planted, it has to stay silent on the repository's own guard code, which contains the same patterns for legitimate reasons, and it has to fail when its input is missing rather than pass on nothing. Exclusions get the same treatment; an exclusion that has not been tested can hide the same problem as a disabled rule.

Deliverables: Gate record

  • Each gate with its three test results
  • The tuned ignore list, every entry attributed
  • Noise rate, factory settings against tuned

Who says no

A gate is ready when the expected test cases pass, including the case where required input is missing. Engineers then fix the finding or document why the risk is being accepted. Once the pipeline gate is enabled, it cannot be bypassed.

04 · Bind

Require the checks before merge.

A security job can run on every pull request and still protect nothing if the branch rules do not require it to pass.

What runs

  • The ruleset. Security jobs become required checks on the default branch, and nobody is exempt from them. This is usually the largest gap.
  • Owners over the agent surface. A code-owners file routes review for instruction files, hooks, plugin configuration, and the gates' own source.
  • Unconditional jobs. A job that runs only when someone applies a label gets skipped by omission. Every security job runs on every pull request.
  • The finding loop. A security finding closes with the fix, and with the instruction that stops the agent regenerating it.

Deliverables: Enforcement map

  • The applied ruleset and its required checks
  • Code-owners coverage of the agent surface and the gates
  • The finding-to-instruction rule, in the instruction files

Who says no

The administrator applies the ruleset.

05 · Re-audit and port

Run the audit again before expanding the rollout.

We compare the result with the original baseline and connect each change to the control responsible for it. If the gates improved the result, we adapt them for the remaining repositories.

What runs

  • The re-audit. The same instrument as the baseline, every delta traced to the control behind it.
  • The port. Remaining repositories in impact-over-effort order, one policy with per-stack implementations held to the same output.
  • The cadence. Scheduled audits, baselines raised as controls land, major tool upgrades reviewed by a person. A scheduler does not notice when an upgrade quietly turns a check off.
  • Bugs in the audit tools. Where the audit itself was wrong, we report the bug to the tool's maintainers with the file, the line and a reproducer.

Deliverables: Posture record

  • Audit delta, control by control, against the locked scorecard
  • Port order with per-repository gate records
  • The review cadence

Who says no

We scale to the rest of the estate only when the head-to-head favors the gates.

The business case

What the work gives you.

You leave with a measured baseline, plus a record of what changed and proof that the required checks now block a merge. That record is what an auditor asks for. It also answers a customer's security questionnaire and gives your own risk discussions something concrete to start from.

One repository · same audit

  • Prevention coverage

    66% 81.5%

  • Supply chain

    83% 96%

  • Application security

    66% 80%

  • Plugin servers pinned

    0 of 4 4 of 4

  • Secret-scan check

    FAIL PASS

  1. 01

    A leak count instead of a guess.

    One scan of your own history gives you a number per repository and a rotation list with owners. In one estate the ungated repository held seven real credentials, one still readable in current code. Its gated sibling was clean.

  2. 02

    Stop secrets before merge.

    A credential remained in one repository for sixteen months. Removing the line did not revoke the credential or erase it from history. The workflow now blocks the write at the agent and commit stages, with the history scan as the final check before merge.

  3. 03

    The workflow covers risks introduced by the agent toolchain.

    Conventional code scanners do not usually inspect whether a suggested package exists, whether a plugin server updates itself, or what commands are hidden in an agent's settings. In one organization, four of six hooks existed only as command strings in a settings file. We include those configuration paths in the review and gate changes to them.

  4. 04

    Every gate proves it fires.

    A control can be present and protect nothing. In one organization security jobs ran and the default branch required none of them to pass, three of four quality gates ran only on a label, and an AI review tool hit its seat limit and went quiet. So each gate here runs against a planted positive, a negative control and an empty input before it is called done.

  5. 05

    The default branch requires the checks to pass.

    In one organization, only one of twenty-six repositories required a passing check before merge; the security jobs in the other twenty-five ran, and the merge went ahead whatever they reported. The audit records whether a check exists and whether it is enforced as separate results. We then use branch rules to make the checks mandatory and review them again as the toolchain changes.

The toolkit

The tools behind the work

The toolkit scans repository history and agent configuration, adds checks to the development workflow, and monitors whether those checks remain enforced.

Layer 01

Read the estate

Establish what has leaked, what is enforced, and what the tooling misses.

Full-history secret scanning

A pinned scanner (gitleaks or TruffleHog) and a single ruleset, run across every repository.

  • Findings split into still present and history only
  • Sorted by credential format, never by testing a credential
  • Noise rate measured at factory settings first

Recurring readiness audit

A dozen scored dimensions, with the evidence behind each check.

  • Application security, supply chain, AI security, prevention coverage
  • Surprising verdicts re-verified by hand
  • A locked baseline the re-audit is measured against

Organization-wide posture scoring

OpenSSF Scorecard across the estate: required checks, dependency updates, branch protection.

  • Read per check, never on the aggregate score
  • Required status checks separated from present ones
  • Dependency update coverage per repository

Agent-surface inventory

Hooks, plugin servers, instruction files and permission patterns, read from the settings files.

  • Hooks defined as inline command strings included
  • Plugin servers and the version each one is pinned to
  • Permission patterns read for credential material

Layer 02

Gate the change

Put a check at every point a change can move, and prove each one fires.

Secret gate

A commit hook, an agent-time guard, and a full-history scan on every pull request.

  • Every ignore-list entry attributed to a named false alarm
  • One value excused at a time, never a whole file
  • Recorded test data kept in scope

Code gate

Static analysis on every merge request (SonarQube, CodeQL), plus a review on the diff before commit.

  • The review runs in a fresh context
  • Confirmed findings written back into the agent instructions
  • A required step. Its catch rate is not yet measured

Dependency gate

Every manifest change checked before the package reaches a developer's machine.

  • The package must exist and pre-date the pull request
  • A 24-hour waiting period on brand new releases
  • Vulnerability audit (OSV-Scanner, Snyk) that no missing label can skip

Toolchain gate

The agent's own configuration treated as security-sensitive code.

  • Plugin servers pinned, and the pin re-verified in the pipeline
  • Hook content scanned, inline commands included
  • A named reviewer through code owners

Layer 03

Keep it true

Make the gates binding, and catch the day one of them stops working.

Rulesets

Required status checks on the default branch, with an empty bypass list.

  • GitHub rulesets, or GitLab protected branches
  • Administrators included in the requirement
  • Code owners over the agent surface and the gates

Posture baseline

OpenSSF Scorecard compared per check against a baseline committed in the repository.

  • The baseline only moves up
  • A dropped check fails the pull request
  • A check that cannot be read fails it too

Accepted-risk register

Deliberate decisions, written down with the reasoning behind them.

  • One named owner per entry
  • Audits stop re-flagging what was already decided
  • Reviewed on the same cadence as the audit

Review cadence

Scheduled audits, and a person reading every major toolchain upgrade.

  • Baselines raised as controls land
  • Secrets migrated to a managed vault with rotation, as a next step
  • A software bill of materials as the next inventory to add

Common questions

What security teams ask first.

We already run a secret scanner. What is different?

We check whether the scanner covers the full repository history, how narrowly it suppresses false positives, and whether a failed scan can block a merge. A scanner that only reads the current working tree can miss a deleted secret that remains in history. When a false positive is silenced by excluding the whole file, any real token later added to that file is silenced with it. And if the branch rules do not require the scan to pass, the finding can still be merged.

We have static analysis. Why is AI-written code a separate problem?

Static analysis still has a place, but it does not cover the whole path an agent can use to change a project. Credentials may appear in recorded test data, generated reports, or agent configuration. An agent may suggest a package that does not exist, while a plugin server can run with the developer's privileges and update outside the normal dependency process. We keep the existing code checks and add coverage for the agent, its dependencies, and its toolchain.

Will the gates slow the team down?

Poorly tuned checks slow a team down and are soon ignored or disabled. Before a rule can block a pull request, we run it against the repository's history and tune its false positives. When a check fails, it points to something in the proposed change that the author can fix.

What do we tell the auditor?

You can show the auditor the original scan, the credential rotation record, the repository rules, and the before-and-after audit. Each change in the audit is linked to the control that produced it. We also document the method, including corrections made along the way.

What will this work not do?

This work stops at the merge; it does not cover network, infrastructure, or runtime security. We never test a suspected credential, because that would mean using it, and credential owners handle rotation in the issuing system. We do not rewrite repository history or include secret values in our deliverables. The agent's pre-commit review is installed as a required step, but its catch rate is not yet measured and is reported that way.

Start with a read-only baseline.
We scan the full repository history and document the leaks, required rotations, audit results, and gaps in merge enforcement before proposing any changes.
Discuss a baseline →