The Number I Wrote Down Was Not Evidence

I had a test count I trusted.
It was 245.
Another method returned 269.
I explained the larger number away.
The text counter must have overcounted parameterized tests, I wrote.
Then I opened the result bundle.
It reported 269 distinct test identifiers.
Forty-one of those nodes carried parameter arguments.
The bundle and its test tree agreed.
I could not reconstruct 245.
So 245 had to go.
Being the author did not make the number primary
The wrong number was not copied from a random comment.
I had written it into a durable rule.
That made it feel reviewed.
It also made the number easier to repeat without reopening the evidence.
Authorship created familiarity, not provenance.
When the derivation disappeared, “I remember checking it” was not a recovery method.
A number I cannot re-derive is a claim with a missing source.
The fact that I created both does not close the gap.
The first explanation was attractively technical
The competing count came from text-shaped output.
Parameterized test systems can render one logical declaration in several ways.
It was reasonable to suspect that a grep-based count had measured presentation instead of test identity.
That suspicion may still be useful.
It did not prove 245.
This is the distinction I missed.
Showing that one reader can be wrong does not make another number right.
Discrediting 269 was not evidence for 245.
The smaller number still needed its own derivation.
The result bundle owned the question
The test run produced an .xcresult bundle.
That artifact contains the structured report for the run.
The current local xcresulttool exposes two relevant readers.
xcrun xcresulttool get test-results summary --path <path>.xcresult
The help describes this as “Get test report summary.”
It also exposes the test tree.
xcrun xcresulttool get test-results tests --path <path>.xcresult
That command is described as “Get all tests from test report.”
In the historical bundle, the summary and the tree converged on 269.
The tree contained 269 distinct identifiers.
Forty-one test nodes had Arguments children, the structural sign recorded for parameterized cases in that bundle.
The exact artifact and structural reader answered the identity question more directly than a grep over rendered text.
The correction was not a vote
Two sources saying 269 did not win by majority.
They won by coverage.
The summary read the test report.
The node tree exposed the individual identifiers and argument-bearing nodes.
The old 245 had no surviving command, artifact query, or intermediate output.
Its evidence basis could not be inspected.
This was not “269 seems more likely.”
It was “269 is re-derivable from the artifact; 245 is not.”
That is a stronger and less dramatic reason.
Withdrawing one claim does not certify the other tool
The original note also warned that counting rendered log lines can misrepresent parameterized tests.
Removing 245 did not prove every grep count correct.
It removed the unsupported comparison.
The general warning can stand on independent evidence.
The specific claim that grep's 269 was an overcount could not.
Corrections should be surgical.
Delete what the new evidence disproves or leaves unsupported.
Do not swing to the opposite universal claim just because one example collapsed.
A test count needs more than a value
The number should travel with its derivation.
At minimum, I now record these fields.
artifact: result bundle path or immutable reference
reader: exact command and tool version
scope: selected test plan, target, and filters
identity unit: test identifier, invocation, or parameter case
revision: source revision that produced the run
observed_at: timestamp of the reading
Without the identity unit, “test count” is ambiguous.
One declaration can generate multiple cases.
One UI line can summarize several nodes.
One log can repeat the same identifier in setup, execution, and summary sections.
The count is meaningful only after the unit is named.
Store the command, not only the answer
A durable note that says “245 tests” is cheap to write and expensive to audit.
A note that stores the exact reader is recoverable.
xcrun xcresulttool get test-results summary --path <path>.xcresult
The artifact path may be ephemeral, so a long-lived record also needs an archived artifact or an immutable reference to it.
The command alone cannot resurrect a deleted bundle.
The bundle alone is awkward if nobody remembers which field was counted.
Provenance is the pair.
Dated readings are not suite properties
The corrected 269 belongs to one historical result bundle.
It is not the current count of that suite.
Tests were added and changed after the run.
This article intentionally does not publish a current total.
A number tied to a revision and artifact is a reading.
Removing the date and revision turns it into a false property of the project.
That is how correct historical numbers become incorrect documentation.
Make the next disagreement cheap
The goal is not to prevent every conflicting count.
Different readers often count different units.
The goal is to make the disagreement diagnosable.
When two numbers differ, print sample identifiers from both sides.
Name the unit each reader counts.
Check whether parameterized arguments are separate nodes or metadata on one node.
Preserve the artifact and exact command.
Then a correction is a comparison, not an archaeology project.
The practical rule
A number I wrote is still a claim.
It needs an artifact, reader, scope, identity unit, revision, and timestamp.
An argument against one counter does not prove another count.
Prefer the reader that owns the structure being counted.
Withdraw a number when its derivation cannot be reconstructed.
Treat every suite total as a dated reading, not a permanent project property.
One last thing
Correcting the note was easy.
Admitting that my own number had no recoverable basis was harder.
But a durable rule should spend authority on things we can still show.
Memory can suggest where to look.
It cannot be the final reader.
Get the next post.
If you made it to the end, meet the next post in your inbox or RSS reader.