Generated Artifacts Can Be Canonical

“Do not commit generated files” is one of those rules that becomes true before anyone finishes saying it.

Usually, it is right.

Build output is reproducible. It is noisy in diffs. It invites merge conflicts. It makes a repository look like someone spilled a dist/ directory on the floor and quietly walked away.

Then a project has a host that serves a generated directory directly. The generated directory is what a push deploys. Suddenly, refusing to commit the artifact does not make the repository cleaner. It removes the thing the deployment contract expects.

The rule did not fail. We asked it a different question.

Ask what role the file plays

The useful classification is not “source or generated?”

It is “what role does this file play in the product?”

An output file can be one of several things:

  • a disposable local cache;
  • a reproducible build byproduct;
  • a release artifact consumed by a host;
  • a reviewed, curated data snapshot;
  • a temporary capture that must stay out of version control.

Those roles deserve different rules.

In the static deployment pattern recorded for one project, the host serves the built output directory with no separate build step in the hosting path. The pushed artifact is the deployment. Committing it is not an accidental exception to a clean repository policy. It is how the product becomes visible.

The important word is intentional. If the output is committed, the repository should say why. Otherwise the next well-meaning cleanup removes a production mechanism with the confidence of someone deleting “unused” kitchen ingredients halfway through dinner.

Curated data is not the same as raw capture

The same distinction applies to a local database.

A database file often looks generated because a program writes it. That does not tell us whether it is disposable.

If a review workflow has deduplicated, selected, and annotated the records in that database, it can be the canonical curated input to a static keepsake. The raw capture files, by contrast, may be noisy, privacy-sensitive, or simply too large to treat as the product’s source of truth.

The line is not “binary files are bad.” The line is “can a reviewer explain why this particular file is here?”

For a curated data artifact, the answer might be:

This is the reviewed record that feeds the published site.

For a raw capture, the answer might be:

This was an intermediate acquisition and is not part of the product record.

Both answers can be correct. Mixing them is where cleanup becomes data loss.

Make the deployment path boringly explicit

An artifact-backed deployment needs three things written down.

First, name the artifact that the host consumes. Do not leave a reader to infer it from a platform dashboard.

Second, name the command or process that regenerates it. If regeneration needs an external capture, a curation step, or a non-obvious environment, say so.

Third, draw the boundary around inputs that should remain ignored. Raw material, secrets, and temporary caches do not become safe merely because a nearby output directory is committed.

That documentation is more valuable than another generic ignore rule. It tells future maintainers which files are deliberately versioned and which ones are merely passing through.

The exception should stay narrow

There is no need to turn this into a philosophy of committing every build folder.

Most applications should still build in CI or on the hosting platform from clean source. Most caches should remain untracked. Most generated files should not appear in a code review.

The exception becomes safe when it is shaped by a concrete consumer:

reviewed input → explicit generator → committed release artifact → host serves artifact

When that chain is real, the generated artifact is not clutter. It is a release record.

The next time someone proposes deleting a generated file, ask one question before reaching for .gitignore:

Is this file merely produced, or is it the product we ship?

That answer should decide the policy.