At the end of September 2026, our team shipped 26 release packages in four days. Counting Free and Pro separately, that was 16 release announcements covering community, jobs, events, learning, gamification, directory and theme products. A small team and a set of AI agents did the work. People who hear that number usually ask the same question: how do you not break things? Our answer is release gates, and this guide shows how our release gates work.

The honest answer is that we do break things, and the interesting part is where. This post is about the system that sits between “the agent says it is done” and “customers get the update”. We call those checks release gates, and every WordPress plugin we ship passes through them. We will show what the gates caught before customers saw it, what slipped past them, and what we changed afterwards. The four-day count is just the example. The lasting idea is how gates work for a small team that uses AI agents to write and test code.

Business owners can read the first half and skip the developer section in the middle. Developers get the actual scripts and checks in a clearly marked part.

First, what release gates are

A release gate is a check that must pass before a version can be tagged and shipped. It is not a reminder, a checklist someone may forget, or an instruction in a chat. It is a step that stops the release when it fails.

The difference matters most when AI agents do part of the work. An agent follows the instructions you give it most of the time, not every time. We wrote about this in why written rules do not reliably govern AI agents. The short version: a rule in a prompt is a request. A gate in the build is a fact. So we move every rule we care about out of the prompt and into something that fails loudly.

That is why a four-day burst of releases is possible at all. Speed comes from agents drafting code, running checks and preparing notes. Safety comes from the gates, which do not care how confident the author sounds.

What a release has to pass

Here is the path a release takes, in plain English. Each step exists in our repositories today.

1. Code quality and static checks

Every commit runs a single check script before it is accepted. It lints the code, runs the WordPress Coding Standards, runs PHPStan static analysis and runs a design-system audit on the interface code. A pre-commit hook runs it automatically, so nobody has to remember. If it fails, the commit does not land.

2. Security sweep

Before a release is tagged, we run a security review across the Free and Pro plugins together. We keep a written list of security gates. It covers injection risks, who is allowed to read or change each record through the API, whether private data leaks into public responses, whether member data is erased and exported properly, and whether secrets are stored safely. The results are written down for that release, with a review of the code changes and a walk through who can call what.

3. The API contract

Our community platform exposes close to 300 REST routes across Free and Pro. A script checks that the committed API description matches the live site: every route is present, every response has a described shape, and the real responses match the description. If a route exists that the documentation does not know about, the check fails.

4. Documentation truth

We keep a machine-readable manifest of every hook, route and setting. A gate fails the release if the manifest has fallen behind the code. This sounds like paperwork, but it protects us from a real problem: an out-of-date manifest tells the next person (or the next agent) that a feature does not exist, so they build it a second time.

5. Behaviour checks

For the larger plugins, a release battery runs behavioural checks against a live test site: a certification run, a suite of user journeys for each role (member, moderator, admin, logged-out visitor), and a static flow audit across the Free and Pro pair. The release script refuses to build the package if these fail. Skipping any of them has to be typed on purpose.

6. Smoke test on a clean install and on an upgrade

A plugin that works on our long-lived test site can still fail on a fresh install, or on a site upgrading from last month’s version. So we install the release zip on a clean container, and we also upgrade from the previous version. We write down what we saw, including the soft findings.

7. A clean package

The release zip is built from an allowlist: only the folders the plugin needs to run are copied in. Notes, screenshots, planning files and development tools cannot slip into the package, whatever happens to be committed.

What the gates caught

These are real catches from the same release window. We list them because they show what kind of problem a gate finds that a code review often does not.

A privacy leak in the “For You” feed

While preparing BuddyNext 1.2.2, we found that the “For You” feed showed posts from spaces a member had joined without checking each post’s own audience. A post set to followers-only or connections-only, made inside an open space, appeared in the For You feed of every member of that space. The space page itself and the single-post API hid it correctly, so the rule depended on which screen you looked at.

We reproduced it on a live test site before changing anything: a followers-only post by one user showed up for another member who was not a follower. The fix applies the shared audience rule to that feed. The test that now guards it walks the For You feed as one of its surfaces, and we confirmed that reverting the fix makes the test fail by name. The fix shipped in 1.2.2, before any customer could see a post they should not.

Two public routes leaking hidden items

In a security sweep of BuddyNext Pro in late September, we found two public read routes that returned objects which were meant to stay hidden. Both are now gated. The same sweep found that provider secrets were stored without encryption in the database, so we now encrypt them at rest and mask the field in the admin screen.

Public routes are a classic blind spot. A route marked public is easy to treat as “anyone can see everything here”, when the correct rule is “anyone can see what they are allowed to see”. Our written security gates now say so directly, including for list totals and counts, which can leak the existence of hidden rows even when the rows themselves are filtered.

The same class of problem in other plugins, in the same week

The security sweeps also produced fixes in several other plugins that shipped in that window. These come from the release notes of each plugin:

  • A private-list plugin: the members of a private list could be read by anyone, including logged-out visitors, and any signed-in user could gain access to a private list by following it. Both are fixed, and following now requires permission to see the list.
  • A business directory plugin: a member could change another business’s team roles or take over a business they did not manage. Fixed.
  • A learning platform: a signed-in member could open another member’s order confirmation by changing the link. Now the confirmation is visible only to the buyer, and guest checkout is rate-limited.
  • A roadmap plugin: several REST routes checked a site-wide capability instead of permission for the specific post, so a Contributor could edit or delete content they did not own. Capability checks are now per post.
  • A My Account customisation plugin: the “visible to roles” setting hid the menu item but did not block the page when someone opened its address directly. It now blocks the page too.

Notice the pattern. Almost none of these are exotic. They are all some form of “the interface hides it, but the server still allows it”. That is exactly what a systematic authorisation walk finds, and exactly what a quick visual check misses.

A scanner that suddenly complained about 98 templates

One catch was a judgement call. During the BuddyNext 1.2.3 smoke pass, the template-contract scan, which checks that templates follow the plugin’s conventions, jumped from 5 failures to 98. The templates it flagged were ones that release had not touched.

A gate that fails is not the same as a bug. We treated this as drift in the scanner’s results rather than a regression in the product, because the flagged files were unchanged, and we kept the previous baseline and recorded the reason in the commit message. This is the kind of decision a person makes, not an agent. If the call had been wrong, the cost would have been a layout contract slipping, so we wrote down exactly what we compared. We would rather have a documented judgement call than a gate quietly switched off.

What still got through

This is the part most “how we ship fast” posts leave out. Our gates missed things too.

BuddyNext 1.2.3 had to follow 1.2.2 the next day

BuddyNext 1.2.2 shipped on 30 September 2026. BuddyNext 1.2.3 followed on 1 October. The 1.2.3 release notes list what the first release had missed:

  • Members and Spaces directory pages 2 and beyond showed “Page not found” instead of the next page of results.
  • The invite-link endpoints answered differently for a secret space the member could not see, which revealed that the space existed.
  • The Members directory printed a stray line of code under the member grid, and a hook that should have fired did not.
  • “Log in” links on the guest banner and in a few other places did not return people to the page they were on.
  • On block themes such as Twenty Twenty-Five, community pages did not use the theme’s own header and footer, and logged a deprecation notice on every load.

None of our gates caught these before 1.2.2 went out. What they share is that each depends on a situation a default test pass does not visit: page 2 of a directory, a space the viewer is not allowed to know exists, a site running a block theme, a guest who clicks Log in from a lightbox. Our checks had measured the common path well and these edges not at all.

The 1.2.3 work added a regression test for the secret-space case, so that exact leak now fails by name. We also ran the clean-zip install and an upgrade from 1.2.2 for the hotfix itself. The wider lesson is that gates only protect what they measure, so each miss should become a new measurement.

The file name that hosting firewalls reject

Two of our plugins had an admin template file named shell.php. Several hosting firewalls block zip uploads that contain a file with that name, because it matches the name of a common malware file. The file was harmless. The install failed anyway on those hosts. We renamed it to layout.php in both plugins.

It is a good example of a gate gap that has nothing to do with code quality. Our checks tested whether the code was correct, not whether a firewall would accept the zip. A filename scan on the final package is cheap, and we recommend adding one to any release process.

A green tick that proved less than it looked

Our own developer notes record one more weakness. The main check script silently skips the journey suite and the certification run unless two environment settings are present. That means a green result from it proves less than it appears to, unless you export the settings when you use it as a release gate. We know about it, we documented it, and the release script now treats those behavioural gates as required instead of optional. But it is a reminder that a gate can pass because it did not run.

Where the AI agents fit

We use AI agents for most of the mechanical work in this process, and it is worth being exact about which parts.

  • Agents write and fix code against a written brief, and they open the branch for review.
  • Agents run the gates and report the results, including the output of the failing check.
  • Agents draft the release notes in a fixed format (New, Improve, Fix, Security, Dev, Compat), one plain sentence per line.
  • People decide what ships. A person reads the security findings, makes the judgement calls (like the template scan above), approves the tag and owns the outcome if a release misses something.

That split is deliberate. We described the broader picture in what a year of AI agents changed in our WordPress agency, and the ownership question in who owns it at 3am. For release work, the practical rule is: an agent can tell you a check passed, but the check has to be something that can fail without the agent’s cooperation.

Why gates beat instructions

A sentence in a prompt cannot stop a build. A release script that exits with an error can. Our own release script carries the history of this lesson in a comment: the certification run and the flow audit used to skip themselves whenever the test site settings were missing, so a machine that had never been set up could package a release with two behavioural gates never running, announced in one line of output that scrolled past. Both are fatal now, and every bypass has to be typed on purpose, so a skipped check shows up in the command history instead of hiding inside a summary.

For developers: the checks behind the scenes

This section is for the engineers who want the specifics. Skip it if you are not technical.

Per commit

  • A pre-commit hook runs the repository check script on staged files. The script runs PHP lint, PHP_CodeSniffer with the WordPress Coding Standards, PHPStan and a UX audit. It is enabled per clone with git config core.hooksPath .githooks.
  • Smaller scripts cover specific rules: REST boundary (frontend talks to REST, not admin-ajax), icon set, dialog patterns, tap-target size, option defaults, store collisions, public hook documentation and erasure coverage.

Per release

  • Security: a written set of numbered security gates with an issue catalogue and a sweep script. Gate 2 is the authorisation walk: for every member-scoped route, confirm the handler checks the specific resource, and for every public route confirm the handler still checks per-object visibility for the current viewer. List routes must exclude hidden rows from totals, counts and returned id arrays.
  • OpenAPI: a script generates a fresh spec and fails if any operation’s success response has no body schema, if the committed spec differs from the fresh one, or if sampled live responses drift from the schema. It compares sampled rows as a union of keys, with a marker for fields that appear only conditionally.
  • Documentation truth: the build script compares the manifest against the code and fails if the manifest is stale. There is an override for a genuinely surface-neutral commit, which has to be typed and explained in the commit message.
  • Behaviour: the release script runs the certification, the journey suite, the flow audit and the unit tests. Each has a named bypass variable, and bypasses are meant to be typed deliberately.
  • Package: the zip is built from an allowlist of runtime folders. Optional folders such as languages and the readme are copied only if present.

Per regression

Every miss gets a test that fails when the fix is reverted. For the For You leak, the existing audience test was extended so that reverting the fix fails by name with a message naming the leaked post. For the secret-space case, a test asserts that the invite-link endpoint returns “not found” for a space the member cannot see. If you cannot make a test fail by removing your fix, the test is not guarding anything.

What another small team can copy

You do not need our stack to use the idea. Start with these.

  1. Put the rule in the build, not the prompt. If your release process depends on an agent or a person remembering, it will fail on the busiest day.
  2. Make skipping a check loud. If a gate can skip itself because an environment variable is missing, it will, and nobody will notice. Make the missing variable an error.
  3. Sweep authorisation as a separate pass. Interface checks tell you what a user sees. Authorisation checks tell you what the server allows. Most of the security fixes above were the gap between the two.
  4. Test an upgrade, not just an install. A fresh install is the easy case. Customers mostly upgrade.
  5. Inspect the zip. Check what is in the package and what a hosting firewall will make of it.
  6. Turn every miss into a measurement. After a hotfix, ask which check would have caught it and add that check, with a test that fails when the fix is removed.
  7. Keep a person accountable. An agent can run the gates. Somebody still has to decide to ship and own what happens after.

Our earlier piece, five releases in a week, and almost none of it was features, looked at the same rhythm from the customer feedback side. This post is the engineering side of it.

If you want to start with the free tooling behind the first gate, the main pieces are the WordPress Coding Standards for PHP_CodeSniffer, PHPStan and the Plugin Check plugin. They take an afternoon to wire into a commit hook, and they catch the boring problems before a person has to.

Common questions

Is shipping this fast safe?

Speed alone is not safe, and the 1.2.3 hotfix shows why. What makes it workable is that the checks run automatically, a miss becomes a new check, and a person owns each decision. We would call it fast with guardrails, not fast and careless.

Do AI agents ship releases without human approval?

No. Agents write, test and draft. A person reviews the security findings and approves the tag.

Do you run the same gates on every plugin?

The principle is the same everywhere, and the depth scales with the product. The community platform has the longest battery because it holds member data at scale. Smaller plugins run the static checks, a security review, the documentation check and a smoke test.

What is the single most useful gate?

If we had to keep one, it would be the authorisation walk. A hidden button that still works on the server is the most common serious bug we find, and no amount of visual testing finds it.

How do you handle a gate that fails for a reason that is not a bug?

You write down why. The template scan in this post is the example: we recorded what we compared and why we kept the baseline. A gate you silence without a reason is a gate you have lost.

The takeaway

Twenty-six packages in four days is not a flex. It is what a good set of gates allows, and it comes with an uncomfortable fact: some releases will still need a fast follow-up. The goal is not zero misses. The goal is that the misses are small, the serious problems are caught before customers see them, and each mistake makes the next release safer.

If you run a plugin catalogue or a client platform and want release gates that hold up when agents are doing part of the work, our team builds and operates exactly this way. Tell us what you ship and how often, and we will show you where the first gate belongs.