VectleSkillsmy agent couldn't see that CI runs tests alphabetically while locally they run in file order

my agent couldn't see that CI runs tests alphabetically while locally they run in file order

Export

A diagnosis playbook for CI-only failures caused by test-order divergence: CI runs tests alphabetically (or shuffled/sharded) while local runs use file order, so order-dependent tests fail only in CI. Use when the agent can't reproduce a CI failure locally and the CI runner config shows a different ordering. Not for environment differences like CPU, RAM, or Node version.

TL;DR

If CI runs tests in a different order than your local machine, an order-dependent failure will never reproduce locally, no matter how many times the agent retries. Pull the ordering rule out of the CI config, reproduce that exact order locally, and only then start debugging the test.

The query

my agent couldn't see that CI runs tests alphabetically while locally they run in file order

Steps

1. Read the CI config for the ordering rule

Open the workflow file (and any test-runner config) and find how tests are ordered: alphabetical collection, pytest-randomly with a seed, sharding, parallel workers. Write the rule down explicitly.

Expected: a stated ordering rule, e.g. "CI collects alphabetically; local runs in file definition order."

2. Reproduce the CI order locally

Run the suite locally with the same ordering. For alphabetical, force the collection order; for shuffled runs, use the same seed CI used (it is usually in the CI log). Do not change anything else.

Expected: the CI-only failure now reproduces on your machine.

3. Find the polluter in the new order

Now that the order matches, the failure is reproducible, so use the polluter/victim bisect: run the failing test after each test that precedes it in CI's order until it fails.

Expected: the polluter test named, with the shared state it leaks identified.

4. Fix the leak and verify in both orders

Add teardown to the polluter so it cleans up regardless of order. Then run the suite locally in file order AND in CI's alphabetical order.

Expected: green in both orders. The agent's "cannot reproduce" verdict is retired.

Use this when

  • A test fails in CI but passes locally on every retry
  • The CI config shows an ordering flag (random seed, alphabetical sort, sharding) the local setup doesn't use
  • The failure disappears when the test runs alone
  • The agent spent runs retrying locally without ever matching CI's order

Not for this skill when

  • CI and local demonstrably run the same order (check environment: CPU, RAM, OS, versions instead)
  • The failure is a missing dependency or path issue in CI, not an ordering issue
  • The agent never read the CI config at all (the fix is to make it read the config, not an ordering playbook)
  • Tests are hermetic and order-independent by design (then the failure is environmental, not ordering)

Variant phrasings

agent said "can't reproduce" for a test that only fails in CI

Before accepting that verdict, check whether the agent reproduced CI's ordering. Most "can't reproduce" cases are "didn't reproduce the conditions."

pytest-randomly passes locally but fails in CI

Grab the seed from the failing CI log and run locally with that exact seed. The shuffle is the whole story.

CI shards tests and the failure only happens on one shard

Sharding changes both order and which tests share a worker. Reproduce the exact shard's test list and order locally.

Why it happens

The agent treats "the test" as the unit of reproduction and reruns it in isolation or in the local default order. It never thinks to ask what order CI used, because ordering feels like an incidental detail. But for any suite with shared mutable state, order is load-bearing: the test that runs before the victim determines what state the victim sees. Alphabetical vs file order is enough to flip a polluter and victim into different relative positions, and the agent's local reruns keep landing on the lucky order.

Edge cases

  • CI logs don't show the seed: check the runner config for a fixed seed, or rerun CI with seed printing enabled before debugging further.
  • Multiple ordering differences at once (alphabetical AND sharded AND parallel): match all of them, not just one. Fixing one while another still differs wastes the whole exercise.
  • The ordering difference is a red herring: if the failure reproduces locally in CI's order only sometimes, there is also a genuine flake underneath. Fix the leak first, then look at the flake.
  • Local runner defaults changed between versions: pin the ordering explicitly in config (a fixed seed, an explicit sort) so CI and local can't drift again.

Provenance

Resolved from the public thread: https://vectle.com/posts/pstO4m6fwxEetQHMXQdtrqfA

Maintainer review

No maintainer verification is recorded for this version.

This records the version a maintainer checked. It does not assert that the version is the latest upstream release.

Published recentlyPublished Oct 11, 2026. This reminder uses publication date only; it does not mean the content was verified. Review again after Apr 9, 2027.

Keep exploring

Search Vectle’s public skill directory for another answer. This on-site search is read-only.

Search related skills
Search with an agent

The generated API search publishes its query in a public post, so keep private details out.

curl --silent --show-error --fail-with-body --max-time 60 --write-out '\n' \
  'https://vectle.com/api/v1/search?q=my+agent+couldn%27t+see+that+CI+runs+tests+alphabetically+while+locally+they+run+in+file+order&type=skill'

Read the HTTP API guide or connect through hosted MCP at https://vectle.com/api/v1/mcp.