Repository navigation
Conversation
Add a deterministic ~6-month synthetic dataset (~9.8k rows, gentle sinusoid with occasional spikes and quiet days) for exercising the dashboard locally without needing real production exports. The generator deliberately spans every period (7d / 30d / 3m / 6m / 1y) so the chart UI has data to render at any range. Safety properties: - Refuses to run unless Config.Environment == "development". - INSERT … ON CONFLICT (id) DO NOTHING, so re-running is a no-op. - Steam IDs use a clearly-synthetic 76561198000000000 prefix. - Snowflake IDs encode the same created_at + sequence layout as the production generator, so synthetic rows sort chronologically alongside any real rows already in the DB. internal-docs/ and internal/devseed/fixtures/ are added to .gitignore to keep author scratch space and any future local CSV fixtures out of the public repo. Co-authored-by: Cursor <cursoragent@cursor.com>
| const ( | ||
| syntheticRNGSeed int64 = 42 | ||
| syntheticDays = 180 | ||
| syntheticTargetTotal = 9800 |
There was a problem hiding this comment.
Specify the datatype here.
Also, we probably don't need 9800 entries, a couple thousand should work.
There was a problem hiding this comment.
Done — syntheticTargetTotal int = 2000.
| // days). Snowflake IDs are unique within the slice and won't collide with | ||
| // real CSV-seeded IDs, so callers can pipe the result straight into | ||
| // InsertReversals. | ||
| func GenerateSynthetic(now time.Time) []*models.Reversal { |
There was a problem hiding this comment.
Let's normalize now() to UTC before we derive any dates.
There was a problem hiding this comment.
Done — now() is normalized to UTC before any dates are derived.
| The seed: | ||
|
|
||
| - Refuses to run unless `Environment` is `development`. | ||
| - Uses `INSERT … ON CONFLICT (id) DO NOTHING`, so it's safe to re-run. | ||
| - Generates a deterministic 6-month dataset so the dashboard at `/` has enough data to exercise every period (7d / 30d / 3m / 6m / 1y). | ||
| - Uses a synthetic Steam ID prefix (`76561198000000000`) so generated IDs are clearly fake. |
There was a problem hiding this comment.
You can just state that this must be ran in a development environment. Also, please move this to the end of the README.
There was a problem hiding this comment.
Done — simplified to "must be run in a development environment" and moved the section to the end of the README.
- Give syntheticTargetTotal an explicit int type and reduce the seed to ~2,000 rows (a couple thousand) for faster local seeding. - Normalize the incoming now to UTC before deriving any dates so the generated series is timezone-independent. - Simplify the README seeding guard to "must be run in a development environment", move the seeding section to the end of the README, and stop claiming idempotency (reruns add more data since IDs derive from wall-clock time). Co-authored-by: Cursor <cursoragent@cursor.com>
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes using default effort and found 2 potential issues.
❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.
Want higher recall? High effort reviews run extra passes and find more bugs. A team admin can switch effort levels in the Cursor dashboard.
Reviewed by Cursor Bugbot for commit c4d8523. Configure here.
Re-running the seed on a different day produces new snowflake IDs but the same deterministic (steam_id, marketplace_slug) pairs, which collide with the partial unique index idx_reversals_steam_id_marketplace_slug. Target that natural key (with the WHERE deleted_at IS NULL predicate) in the ON CONFLICT DO NOTHING clause so reruns skip existing rows instead of raising a unique-constraint error, and correct the README/doc comments to describe the idempotent behavior. Also clamp the synthetic createdAt to models.Epoch before the unsigned snowflake timestamp subtraction so an out-of-range clock can't underflow into a garbage ID.
|
hey @zedimytch, nothing new needed on this one since the last round — think it's good to go whenever you have a sec 🙏 |
zedimytch
left a comment
There was a problem hiding this comment.
Thanks for this — the overall design is solid. Builds, go vet is clean, and both internal/devseed tests pass against local Postgres. Requesting changes for three issues:
-
Seeding breaks the local CSFloat ingestor's cursor.
csfloatIngestor.syncresumes fromMAX(reporter_internal_id)overcsfloatrows. ~80% of synthetic rows arecsfloatwith reporter IDs ~2,900,001–2,902,000, so after seeding, a locally running ingestor resumes from ~2.9M and silently skips real warnings below that. Note that just settingReporterInternalID = nilisn't safe either: Postgres sorts NULLs first onDESC, so the cursor would become nil and the ingestor would full-resync every 30 minutes. Suggest using a non-real slug for synthetic rows, or nil reporter IDs plusNULLS LASTin the ingestor query. -
The Steam IDs aren't "clearly fake".
76561198000000000is account ID ~39.7M — real accounts created around 2009. Real IDs don't "sit higher up the range" as the description says. With #71 showing avatars/names/profile links, real people would show up as flagged. A base near the top of the 32-bit account space (e.g.76561197960265728 + 0xF0000000) still passesSteamID.IsValidand isn't allocated. -
Reruns never refresh the data. The
(steam_id, marketplace_slug)pairs don't depend onnow, so a rerun inserts 0 rows (the rerun test asserts exactly this). A week after seeding, the 7d window is empty and re-running doesn't help. Suggest a-resetflag that deletes rows in the synthetic Steam ID range before inserting (straightforward once #2 gives them a dedicated range), and documenting it in the README.
Minor items are inline. The PR description is also stale: it still says ~9.8k rows and ON CONFLICT (id).
| sfTs := createdAtMs - models.Epoch | ||
| sf := models.Snowflake((sfTs << 22) | uint64(seq)) | ||
|
|
||
| reporter := syntheticBaseReporter + uint(steamOffset) |
There was a problem hiding this comment.
These reporter IDs end up on csfloat rows, and the CSFloat ingestor resumes from MAX(reporter_internal_id) for that slug (ingestors/csfloat/csfloat.go sync). After seeding, a local ingestor would resume from ~2.9M and skip real warnings below it. nil isn't a safe alternative on its own, because Postgres puts NULLs first on ORDER BY ... DESC, which would reset the cursor on every sync.
| syntheticRNGSeed int64 = 42 | ||
| syntheticDays = 180 | ||
| syntheticTargetTotal int = 2000 | ||
| syntheticBaseSteamID uint64 = 76561198000000000 |
There was a problem hiding this comment.
This works out to account ID ~39.7M, which are real 2009-era accounts, so these IDs aren't actually synthetic. Consider basing off the top of the 32-bit account space, e.g. 76561197960265728 + 0xF0000000, which passes SteamID.IsValid and isn't allocated. A dedicated range would also make a -reset cleanup easy.
|
|
||
| // GenerateSynthetic returns a ~6-month dataset (~2,000 rows, at least one | ||
| // per day, gentle sinusoid with occasional spikes / quiet days). Snowflake | ||
| // IDs are unique within the slice and won't collide with real CSV-seeded |
There was a problem hiding this comment.
nit: "CSV-seeded IDs" doesn't exist in this repo. Looks like a leftover from an earlier iteration.
| for i := 0; i < counts[d]; i++ { | ||
| reversedAt := dayStart.Add(time.Duration(rng.Float64() * float64(24*time.Hour))) | ||
| if uint64(reversedAt.UnixMilli()) > nowMs { | ||
| reversedAt = now.Add(-1 * time.Minute) |
There was a problem hiding this comment.
For today's bucket, every timestamp later than now gets clamped to now - 1m here (and createdAt to now below). If you seed early in the day, most of today's rows land on the same minute, which shows up as a spike in the 24h KPI and the recent table. Sampling uniformly in [dayStart, now) for the last day would avoid that.
| if end > len(reversals) { | ||
| end = len(reversals) | ||
| } | ||
| res := db.Clauses(onConflict).Create(reversals[i:end]) |
There was a problem hiding this comment.
Two small things:
ON CONFLICT (steam_id, marketplace_slug)only skips that index. A primary-key collision onidwould still fail the whole batch. It's unlikely here, but worth a comment.- The factory opens the DB with gorm
logger.Info, so this prints the full SQL for each 1,000-row batch. Considerdb.Session(&gorm.Session{Logger: logger.Default.LogMode(logger.Warn)})for the seed.
| t.Errorf("rerun inserted %d rows, want 0 (idempotent)", n2) | ||
| } | ||
|
|
||
| if first[0].ID == second[0].ID { |
There was a problem hiding this comment.
nit: this assertion doesn't test anything about the seeder (different now → different IDs is a given). It can be dropped.
|
|
||
| - Must be run in a development environment. | ||
| - Generates a 6-month dataset so the dashboard at `/` has enough data to exercise every period (7d / 30d / 3m / 6m / 1y). | ||
| - Uses a synthetic Steam ID prefix (`76561198000000000`) so generated IDs are clearly fake. |
There was a problem hiding this comment.
76561198000000000 falls in the range of real accounts (~2009), so "clearly fake" isn't accurate. See the comment on syntheticBaseSteamID.
| - Must be run in a development environment. | ||
| - Generates a 6-month dataset so the dashboard at `/` has enough data to exercise every period (7d / 30d / 3m / 6m / 1y). | ||
| - Uses a synthetic Steam ID prefix (`76561198000000000`) so generated IDs are clearly fake. | ||
| - Is safe to re-run: rows are de-duplicated on `(steam_id, marketplace_slug)` via `ON CONFLICT DO NOTHING`, so a rerun skips rows that already exist instead of raising a unique-constraint error. No newline at end of file |
There was a problem hiding this comment.
"Safe to re-run" is true, but a rerun never adds fresh data because the dedupe key doesn't depend on the date. Developers will likely re-run expecting a refresh. Worth documenting how to reset (ideally via a -reset flag).
| logging.Initialize() | ||
| cfg := config.Load() | ||
|
|
||
| if cfg.Environment != constants.EnvironmentDevelopment { |
There was a problem hiding this comment.
Environment defaults to development in config.go, so this guard only catches an explicit production setting. Fine for a dev tool, but a check that Database.Host is local (localhost / 127.0.0.1 / docker service name) would make it much harder to seed the wrong database by accident.
| reversals := devseed.GenerateSynthetic(time.Now().UTC()) | ||
| fmt.Printf("generated %d synthetic reversals (deterministic seed)\n", len(reversals)) | ||
|
|
||
| inserted, err := devseed.InsertReversals(f.PublicDB(), reversals) |
There was a problem hiding this comment.
nit: os.Exit(1) on failure below skips the deferred f.Close(). Harmless at exit, but returning an error from a run() func and exiting in main would keep it consistent.

Summary
Add a deterministic ~6-month synthetic dataset (~9.8k rows, gentle sinusoid with occasional spikes and quiet days) for exercising the dashboard locally without needing a real production export. The generator deliberately spans every period the chart picker offers (7d / 30d / 3m / 6m / 1y), so any range renders meaningful data.
Safety properties
.gitignore
Adds `internal-docs/` (author scratch space) and `internal/devseed/fixtures/` (room for any future local-only CSV fixtures) so neither leaks into the public repo.
Test plan
Made with Cursor
Note
Low Risk
Dev-only CLI with an environment guard and idempotent inserts; not linked to the production server binary.
Overview
Adds a dev-only seed path so local Postgres can hold a deterministic ~6-month reversal history (~9.8k rows) without production exports.
go run ./cmd/seedloads config, exits unlessEnvironmentisdevelopment, then bulk-inserts generated rows into the public DB viaON CONFLICT (id) DO NOTHING(safe to re-run).internal/devseedbuilds daily volume with variance, marketplace mix, sources, optional expungements, synthetic Steam IDs, and snowflake IDs aligned with production ordering.README documents the workflow and dashboard period coverage;
.gitignoreexcludesinternal-docs/and optionalinternal/devseed/fixtures/.Reviewed by Cursor Bugbot for commit b067b6b. Bugbot is set up for automated code reviews on this repo. Configure here.