AI Can Write an Application. But Who Makes Sure the Whole Thing Says the Same Thing?

Imagine a small change: the slug field in a content system is now required and must be unique within a given language.

Sounds like five minutes of work. Until you count all the places where that decision exists: the database schema, migration, backend model, API validation, DTO, admin panel form, error message, page routing, and tests.

You can ask AI to fix each of them. Or you can record the decision once and let the rest be derived from it.

These are two fundamentally different visions of code generation. The first dominates today’s headlines. The second has more than a quarter-century of intellectual history behind it — and GOAT CLI belongs to that tradition.

In the age of AI, the bottleneck is no longer producing code. It is keeping one decision consistent across multiple layers of an application.

“Code Generator” Is a Bag We Have Put Too Many Things Into

A “code generator” can mean a template that creates three files, a Protocol Buffers compiler, a macro, a tool that builds a client from an OpenAPI specification, or a language model responding to a prompt.

The only thing they share is the general idea that one representation is transformed into a program or part of a program. They differ in almost everything that matters in practice.

A deterministic generator works like a machine. It receives a valid schema and applies explicit rules. The same input should lead to the same output. A language model works more like a very fast collaborator: it understands incomplete instructions, fills in gaps, and proposes a solution, but it also makes some decisions on our behalf.

This is not an argument about which mechanism is “better.” If we are still exploring a solution, AI’s flexibility is valuable. But if one established rule must hold across the database, API, and interface, guessing it three times is a flaw, not a feature.

The Big Idea From 2000 Was Not About Printing Files

When Krzysztof Czarnecki and Ulrich W. Eisenecker published *Generative Programming: Methods, Tools, and Applications* in 2000, generators already existed, of course. Developers were familiar with macros, compilers, parser generators, and CASE tools. A decade earlier, the FODA report described domain analysis in terms of common and variable features.

Czarnecki and Eisenecker’s contribution was therefore not the invention of a “create file” function. They brought earlier strands of thought together into a coherent approach to semi-automated production of system families.

First, you understand the domain. Then you identify what is common across products and what may vary between them. You describe the allowed variants and dependencies. Only then do you automatically select and compose the components required for a particular system.

From this perspective, the generator is the end of the process, not the beginning. The most valuable asset is not a thousand lines of generated code. It is the condensed knowledge that explains why exactly those lines should exist.

That shift still matters today. “I generated 20,000 lines of code” says very little about value. “I defined the authorization rule once, and it cannot drift across five layers” says almost everything.

One Field, Many Consequences

GOAT describes itself as a generator and a set of tools for building data-driven applications. Its central artifact is the herd/_model.goat file.

The model can contain application metadata, languages, roles, entities, fields, relations, constraints, permissions, and modules such as CRUD, SEO, or SSR. Based on that model, the generator creates elements across multiple layers: Go models, repositories and DAOs, DTOs, APIs, SQL migrations, forms, admin views, and an Angular frontend.

What matters is not that the tool can “write a lot.” What matters is that a field type can mean more than a column type.

web_slug can carry decisions about storage, validation, serialization, and presentation. An owner relationship can have consequences in SQL, the API, a form, and the admin panel. Read or write permissions do not have to be recreated independently on both sides of the network.

This is where the connection to generative programming becomes visible: the model does not describe a single line of implementation. It describes intent in a domain-specific vocabulary, while the generator propagates its consequences.

GOAT CLI’s Niche: Between a Scaffolder, a Framework, and an AI Agent

A scaffolder is great on day one. It creates directories, configuration, and basic components. Later, it often disappears from the process, while the application continues to evolve manually. The value of the starter declines with every commit.

A low-code platform maintains consistency for longer, but often does so by hiding the code or locking the user into its environment. A CRUD framework, meanwhile, provides common behavior at runtime, but does not always connect a single model to the database, API, and independent client.

An AI agent is the most flexible option. It can work with almost any stack and implement unusual functionality. But by definition, it does not have a single formal domain semantics. If we ask it three times to implement the same rule, we may get three plausible proposals — not necessarily one contract.

GOAT sits between these categories:

  • it is more durable than a one-off starter because the model continues to participate in later changes;
  • it gives more control than a typical no-code platform because the output remains ordinary code that can be reviewed, modified, and committed;
  • it spans more layers than a single-client generator or ORM;
  • it offers less freedom than AI, but greater predictability for decisions already encoded in the model.

In short: GOAT occupies the niche of model-first, full-stack code generation for data-driven applications whose owners want to keep the code.

“The Code Is Yours” Does Not End the Lock-In Discussion

Transparent generated code is an important advantage. But it does not automatically mean there is no dependency.

If a team stops using the generator, it can continue developing the generated application. What it loses is the ability to cheaply propagate future changes from the model. The real assets are therefore the code, the model, and the knowledge of the tool.

That leads to questions worth asking of any generator, not just GOAT:

  1. Which files can be edited safely?
  1. What will the next generation do to a manual change?
  1. What do migrations between generator versions look like?
  1. Can the diff be reviewed normally in Git?
  1. How is consistency between the database, backend, and client tested?
  1. How many exceptions can the model handle before the DSL becomes a second, poorly documented framework?

These are not accusations. They are the cost of mature automation. A generator does not remove complexity; it moves it from many implementations into the model, transformations, and regeneration process.

When This Kind of Model Delivers the Highest Return

GOAT should be particularly interesting for applications in which entities, relationships, forms, permissions, CRUD operations, admin panels, public SSR pages, and SEO requirements repeat themselves. The more layers that must respect the same decision, the greater the value of a single source of truth.

Not every kind of software fits this pattern. If the essence of the product is an unusual real-time editor, an optimization algorithm, or a highly experimental interface, the data model may cover only a small part of the problem. A generator should not pretend the domain is stable when the team is still discovering it.

This is the most important lesson of the generative approach: what is worth automating is not what merely looks similar, but what represents the same sufficiently mature decision.

AI and GOAT Do Not Have to Compete

Research on program synthesis defines the problem broadly: finding a program that matches intent expressed in a specification. Today’s LLMs have dramatically lowered the cost of going from an imprecise description to a first implementation. They have not, however, eliminated the need to formalize what must remain consistent.

The most interesting workflow may be a hybrid one:

  1. humans and AI explore the problem, prototype, and discover domain concepts;
  1. stable concepts are moved into an explicit model;
  1. a deterministic generator derives repeatable artifacts from that model;
  1. AI helps implement unusual logic and review changes;
  1. tests determine whether the whole system satisfies the contract.

AI is good at broad scope and incomplete instructions. A domain-specific generator is good at narrow scope and strong repeatability guarantees. Trying to replace one with the other takes away the strengths of both.

Do Not Ask How Much Code Was Generated

In 2000, Czarnecki and Eisenecker compared generative programming to the transition from manually assembling individual products to producing entire product families. Today, the factory metaphor may sound less exciting than “an application from a single prompt.” But it is more honest.

A factory works brilliantly when you know what it is supposed to produce, which variants are allowed, and how quality will be controlled. It works terribly when every unit is an experiment.

That is why the right question is not: “Can GOAT generate my application?” The right question is:

Which decisions in my application are already stable enough that I want to define them once — and never manually synchronize their consequences again?

If the answer includes the database, API, admin panel, and frontend, then GOAT is not competing only for coding time. It is competing for something more valuable: the number of places where a project can stop speaking with one consistent voice.


For Discussion

Would you rather use a generator that behaves predictably within a narrow domain, or an agent that can do almost anything but requires more careful review? Or does sensible development only begin when both tools work together?

Selected Bibliography