AI-native WordPress plugin development
Engineering ExMoment Author: Building a Production WordPress Plugin Entirely With AI
How we built and maintained ExMoment Author with zero manual production code authoring, combining Spec-Driven Development, rapid iteration, AI agents and human engineering review.
Project / context
What happens when AI is not only part of the product, but also responsible for writing the product's production code? That was one of the questions behind ExMoment Author, an AI-assisted publishing plugin for WordPress.
From the first implementation through architecture changes, debugging, automated testing and production releases, we developed the plugin using an AI-native engineering workflow with zero manual production code authoring.
- 0
- Manual production code authored in the development covered by this study
- 1.3.7
- Public plugin version checked for this study
- 47
- Passing assertions in the documented taxonomy regression run, not rerun for this study
This was not an experiment in asking an AI to generate a plugin from a single prompt.
Human engineering remained central to the process. We defined the product, requirements, architecture, constraints, expected behavior, acceptance criteria and validation strategy. AI agents inspected the repository, implemented changes, refactored code, investigated failures and produced the production implementation.
The result became ExMoment Author, a real publicly available WordPress plugin that progressed through multiple production releases.
ExMoment Author on WordPress.org provides the public release history. Version 1.3.7 was the current published release checked for this study.
ExMoment Author started from a practical publishing problem. Generating text with an AI model is relatively straightforward. Building a publishing system around that capability is not.
A production WordPress workflow has to understand much more than a prompt and a response. It has to work with posts, titles, HTML content, taxonomies, categories, authors, metadata, source material, editorial rules, validation, model failures and the existing behavior of WordPress itself.
We wanted ExMoment Author to operate inside that environment rather than behave like a disconnected AI text generator. The product therefore evolved into an AI-assisted publishing system in which WordPress remained the application and content-management layer while AI became one component within a controlled generation pipeline.
At the same time, we made another decision that changed the nature of the project:
We would develop the production implementation through AI rather than manually authoring the code ourselves.
Spec-Driven Development and rapid iteration
We combined Spec-Driven Development with Rapid Application Development (RAD). These approaches solved different parts of the problem. Spec-Driven Development gave individual AI implementation cycles structure.
Before asking an agent to modify the software, we increasingly defined:
- objective
- relevant context
- constraints
- expected behavior
- implementation boundaries
- validation requirements
RAD gave us the iteration model.
Instead of attempting to completely define the finished product before implementation began, we worked through short cycles of:
- Plan
- Implement
- Test
- Evaluate
- Refine
- Repeat
Working software continuously informed the next specification.
RAD determined how quickly we iterated. Spec-Driven Development determined how tightly each iteration was controlled.
This combination was particularly useful when developing AI functionality because some requirements only became clear after observing actual model behavior inside the application.
The problem
Throughout the development covered by this case study, production implementation changes were generated and applied through AI-driven development workflows. Human engineers remained responsible for product direction, requirements, architectural guidance, constraints, review, testing strategy and final acceptance.
We did not manually author the production implementation. That does not mean the software built itself. Removing manual code authoring increased the importance of other engineering activities. Requirements had to become more precise.
Expected behavior had to be observable. Architectural boundaries had to be communicated clearly. Acceptance criteria had to be explicit. Failures needed to produce useful information for the next iteration. The developer's role moved away from directly manipulating implementation details and toward controlling intent, architecture and verification.
Constraints / challenges
One rule stayed the same:
The fact that AI produced an implementation did not mean the implementation was accepted. An agent can produce code that is syntactically correct, internally consistent and still wrong for the product. It can misunderstand a requirement.
It can solve the visible symptom instead of the architectural problem. It can make a reasonable assumption that conflicts with existing WordPress behavior. It can produce an elegant implementation of something we did not actually want.
The development loop therefore could not end when the agent reported that a task was complete. Completion required verification against the specification.
Long-running AI-assisted projects introduce another challenge:
Context is temporary. A human engineer working on the same product for months develops an internal mental model of the system. A new AI context does not automatically inherit that understanding. Architectural intent therefore had to become discoverable.
Repository structure, service boundaries, naming, tests, specifications and existing implementation patterns became forms of communication between development sessions. Maintainable architecture was no longer useful only for the next human developer.
In an AI-native workflow, architecture also became context infrastructure for the next agent.
The solution
Early AI-assisted development can feel deceptively simple. Describe a change, allow the model to inspect the code and ask it to implement the requirement. That approach becomes less reliable as software grows. A mature repository contains assumptions that are not visible from a single file.
Features depend on existing services, WordPress hooks, database state, naming conventions, tests and behavior elsewhere in the application. We found that loosely phrased development requests left too much room for interpretation.
The response was not to make prompts longer simply for the sake of length. Development instructions became more structured. That process evolved into our CD3 workflow.
At a conceptual level:
- Objective
- Relevant context
- Constraints
- Implementation requirements
- Validation requirements
- AI execution
- Verification
- PASS / PARTIAL / FAIL
The important principle is:
Implementation should begin from an engineering specification rather than an ambiguous request.
ExMoment Author gradually developed clear boundaries between WordPress, the generation workflow and the AI layer.
Conceptually:
- WordPress
- Generation job
- Source/context preparation
- AI service
- Model/provider layer
- Structured result
- Application validation
- WordPress persistence
- Editorial review
This separation mattered. We did not want the application to become a collection of prompts directly connected to WordPress database operations. The AI layer should generate or propose information. The application should remain responsible for deciding whether that information is valid and what happens to it.
If zero manual coding meant repeatedly asking an AI to change things until the application appeared to work, the process would become increasingly fragile as the codebase grew. Our experience pushed us toward the opposite model.
The less implementation code we manually authored, the more important specifications, architecture, deterministic validation, automated tests and explicit acceptance became. Spec-Driven Development gave agents bounded problems.
RAD allowed us to learn quickly from working implementations. CD3 gave us an increasingly structured mechanism for transferring engineering intent into executable AI development cycles. AI operated inside that system.
Technical implementation
One of the early lessons was that prompts could not carry the entire reliability burden. The generation requirements became increasingly sophisticated.
The system needed to reason about source material, synthesize rather than blindly reproduce it, produce appropriate WordPress content, respect editorial behavior, generate usable titles and return information in formats the application could validate.
We introduced optional context based on the public WordPress author display name and per-job system prompt overrides as the editorial model became more flexible. A job override changes editorial direction while the mandatory output requirements remain active.
The temptation with an AI application is to solve every behavioral problem by modifying the prompt. We increasingly moved in the opposite direction. If something could be enforced deterministically by the application, we preferred deterministic enforcement.
Prompts remained responsible for tasks that benefited from language understanding and generation. PHP and WordPress remained responsible for rules that could be expressed as application logic. "Reliable AI behavior did not come from discovering a magical prompt. It came from reducing ambiguity across the entire system."
One architectural principle that emerged was treating model output as data crossing a trust boundary. A response can look convincing while containing information the application should not accept. This matters because generated output can influence persistent WordPress state.
Taxonomy became a strong example. As categorization evolved, we moved toward constrained category identifiers rather than allowing unrestricted control over taxonomy values. The AI could select from an allowed set of category slugs.
The application validated those values strictly before using them. WordPress hierarchy rules remained deterministic. Where parent relationships were required, application logic resolved and assigned the appropriate ancestors.
Let AI propose. Let deterministic application logic decide what is allowed.
Another important architectural change was reducing direct dependency on a single model integration. ExMoment Author moved toward the WordPress AI Client and a provider-neutral AI service abstraction.
The provider boundary:
- Publishing workflow
- AI service
- WordPress AI Client
- AI provider/model
The publishing workflow should understand the capability it needs, not every provider-specific implementation detail required to obtain it. This created a cleaner boundary for model interaction, validation and future provider changes while keeping the publishing workflow focused on application behavior.
As ExMoment Author evolved, editorial generation became more deliberate. The system distinguished source analysis from content synthesis. Generated content was expected to understand the central idea of its source rather than simply rearrange source paragraphs.
Output requirements became explicit. WordPress-ready HTML was preferable to arbitrary formatting conventions leaking into the editor. Titles developed their own policy rather than simply being extracted from generated body content or excerpts.
Author context and per-job editorial instructions became configurable rather than requiring a separate implementation for every editorial identity. These individually small changes illustrate how an AI prototype becomes an application.
The model remained generative. The surrounding software became increasingly deterministic.
The zero-manual-code constraint changed debugging substantially.
A traditional workflow might be:
- Find bug
- Open file
- Edit implementation
- Test
Our workflow became:
- Observe incorrect behavior
- Reproduce the problem
- Define expected behavior
- Gather repository context
- Provide evidence and constraints to the agent
- Agent investigates
- Agent modifies implementation
- Run validation
- Review result
- Accept or iterate
For extremely small changes, manually editing an obvious line could have been faster. We deliberately avoided that shortcut. The point of the experiment was not merely to determine whether AI could generate an initial codebase.
We wanted to know whether an AI-native development process could survive normal software maintenance.
If humans were not manually authoring production code, how could we trust it? We did not trust code because AI generated it. We trusted behavior after verification. Automated tests became particularly valuable because they provided both humans and agents with an objective feedback mechanism.
During the categorization rework, validation and hierarchy behavior became executable through tests. The categorization audit distributed with the public plugin records a regression run in which all 47 assertions passed. That is a historical validation record, not a claim that this case study reran the suite.
The important part was not the number itself. Expected behavior had become executable. An agent could change implementation details while tests continued to express what the system was expected to do. Tests also made iterative development safer because later changes could be checked against established behavior rather than reconstructed from memory.
Outcome
The strongest test was whether the project could continue evolving after its first successful implementation. ExMoment Author did. The plugin progressed through multiple releases as its AI integration, categorization, editorial behavior, validation and other capabilities evolved.
Release engineering remained part of the AI-native workflow, including version changes, validation, packaging requirements and WordPress.org SVN deployment preparation. The public plugin and release history matter because this was not an isolated proof of concept that stopped once a demonstration worked.
The experiment produced software that went through a real plugin lifecycle and continued evolving as requirements changed.
There is an important distinction between code authoring and engineering. For the development covered here, production source-code changes were generated and applied through AI-assisted and agentic workflows. Humans remained responsible for deciding what should exist and whether the resulting implementation was acceptable.
Human engineering
Human responsibilities included:
- product requirements
- priorities
- architectural direction
- technical constraints
- WordPress behavior
- integration expectations
- security expectations
- acceptance criteria
- testing strategy
- reviewing outcomes
- identifying incorrect behavior
- directing subsequent iterations
- approving releases
AI implementation
AI agents were responsible for:
- repository inspection
- implementation changes
- refactoring
- creating/modifying tests
- investigating failures
- carrying out code-level work required by specifications
- validation execution
- release preparation
We did not remove the engineer from development. We changed where engineering effort was applied.
ExMoment Author demonstrated something more useful than the fact that AI can generate production code. It showed that a software-development workflow can be reorganized around AI implementation while retaining human control over product intent and engineering decisions.
This required clearer specifications, explicit constraints and acceptance criteria. We also had to separate generative behavior from deterministic application logic and give agents enough context to understand the existing system. The resulting process combined rapid iteration with increasingly explicit engineering control. The developer did not disappear. The developer moved.
The work shifted from manually expressing every implementation detail in PHP toward designing the system, specifying its behavior, creating the boundaries within which AI could operate and determining whether the result was actually correct.
What we learned
AI did get things wrong. There were development cycles where an agent misunderstood what we wanted. There were implementations that solved the technical requirement but missed the intended product behavior. There were changes that required architectural correction after their interaction with the wider system became visible.
There were also cases where the generated code was not the primary problem. The specification had left important behavior ambiguous. The agent had insufficient repository context. An acceptance criterion existed in our heads but had not been written down.
A requirement explained what should change without explaining what must remain unchanged. Those failures changed the development process. Specifications became more explicit. Regression requirements became more important.
Agents were instructed to inspect existing implementations before modifying them. Validation became part of the directive rather than an informal activity performed afterward. A human developer familiar with a repository can silently compensate for an incomplete specification.
An AI agent can expose the missing information by making a reasonable but incorrect interpretation. That feedback made our specifications better.
Traditional development spends significant engineering time translating an already-understood solution into source code. AI can compress parts of that translation. The engineering effort does not disappear.
Attention moves toward questions such as:
What exactly should this feature do? What should it never do? Which existing behavior must remain unchanged? Where does this responsibility belong architecturally? What information does the agent need before modifying the implementation?
How will we know that the implementation is correct? What happens when the model returns unexpected data? What regression would tell us another part of the system was damaged? These are software-engineering questions, not prompt tricks.
The experiment began with a straightforward question:
Can we build a WordPress plugin without manually writing its production code? That question became less interesting as the project progressed. AI could write the code. The harder problem was building a development process that remained reliable as the product became more complex.
Specifications became first-class engineering artifacts
When implementation is delegated, ambiguity becomes expensive. A requirement a human developer might resolve through intuition can send an agent toward a completely different implementation. Writing down intent became engineering rather than administrative overhead.
Context became infrastructure
Agents perform better when they understand the repository, architecture, constraints and existing behavior surrounding a requested change. Providing that context consistently became part of the development system.
Tests became a communication mechanism
A good automated test does more than detect regressions. It communicates expected behavior to future implementation agents. Tests became durable context that could survive between AI sessions.
Deterministic rules should remain deterministic
AI is valuable where interpretation, language understanding and generation are required. It should not replace predictable application logic simply because AI is available. Taxonomy validation demonstrated this principle.
The model could make a semantic selection. WordPress and application logic remained responsible for deciding whether that selection was valid and how it should affect persistent state.
Failures improved the development system
Incorrect implementations were useful when they exposed ambiguity in specifications, missing tests or weak architectural boundaries. The goal was not to prevent an agent from ever making a mistake. The goal was to create a workflow where mistakes were observable, diagnosable and capable of improving the next development cycle.
The experiment stopped being about whether AI could write code. We already knew it could. The interesting question became whether we could build a development process around AI that remained reliable as the software became more complex.
ExMoment Author became our practical answer to that question. The public plugin and its release history can be viewed through the WordPress.org reference in Project / context.
Public plugin and release history
Explore ExMoment Author on WordPress.org
Review the public plugin and its release history.