Blogs.BlogId was a one-time copy of BlogNames and nothing kept it current: 12,238 blogs first seen after 2026-08-07 had a BlogNames ID but a NULL Blogs.BlogId, so GetBlogs' join on BlogId silently skipped them and their 23,148 notes. - retire-blognames.sql: stub Blogs rows for the 17 unregistered note participants, backfill IDs (none renumbered), make ix_Blogs_BlogId UNIQUE, drop BlogNames, and add triggers that stop a Blogs row with a BlogId from being deleted, renamed or renumbered - AddNote registers both blogs via RegisterBlog (Blogs row + MAX+1 ID) and every query resolves names through Blogs instead of BlogNames - verify-db-schema.sql reports a DB that still has BlogNames (1e) - Update TL.db.md, AGENTS.md and the DB Browser saved queries Co-Authored-By: Claude Opus 5.5 <[email protected]>
215 lines
14 KiB
Markdown
215 lines
14 KiB
Markdown
# AGENTS.md
|
|
|
|
## Project Overview
|
|
- **Primary Language**: C# (.NET 8 Console Application)
|
|
- **Key Libraries**: RestSharp, Newtonsoft.Json, System.Data.SQLite, Microsoft.Extensions.Configuration
|
|
- **Purpose**: Tumblr API data harvester for collecting notes, posts, likes, and replies, storing results in SQLite.
|
|
|
|
## Architectural Patterns
|
|
- CLI entry point in `Program.cs` with workflow orchestration
|
|
- `DataAccess.cs`: Database operations, `ApiKeyPool` (API key management), `APIAccess` (Tumblr client)
|
|
- `ResponseNotes.cs`: Tumblr API response models
|
|
- Round-robin API key rotation with rate-limit tracking
|
|
- Automatic console color assignment per API key for output differentiation
|
|
- No-argument mode (`Program.TraverseDirectory`) ingests `.txt` blog export files into `Posts` via `DataAccess.AddPost`. Recognized field prefixes live in `TraverseDirectoryFieldPrefixes`; `Body:` and `Downloaded files:` collect every following line up to the next recognized prefix (multi-line values). `RootURL` is populated from a `Reblog root url:` line the same way it's populated from the API-based `--likes` flow — both paths converge on `DataAccess.AddPost`'s `rootURL` parameter, which `UpdatePost` only overwrites when the incoming value is non-empty (existing `RootURL` is preserved otherwise)
|
|
|
|
## Developer Guidelines
|
|
|
|
### Code Formatting
|
|
- 4-space indentation, no tabs, match existing C# style
|
|
- PascalCase for public members, camelCase for locals
|
|
- Minimize code comments unless explicitly requested
|
|
- Use only existing project libraries; no new dependencies without confirmation
|
|
- Match accessibility modifiers (`public` for models, `internal` for helpers)
|
|
|
|
### Error Handling
|
|
- Wrap file/network operations in `try-catch`
|
|
- Log non-critical errors (e.g., config write failures) with `[Warning]` prefix
|
|
- Preserve console color state: use save/restore pattern for temporary color changes
|
|
- API rate limits must use `ApiKeyPool.MarkRateLimited()`/`MarkAvailable()`
|
|
|
|
### API Failure Classification
|
|
Tumblr sits behind a CDN that returns HTML error pages (403, 5xx) which never reach the API. These
|
|
say nothing about the item being fetched, so they must not be recorded as per-item failures.
|
|
|
|
- A response body that will not parse as JSON did not come from the API. Flag it with
|
|
`Root.transientFailure`, never as `FAILURE`
|
|
- Transient failures retry in place (`TransientBackoffSeconds`) before the item is skipped; a skipped
|
|
item stays unmarked in the DB so a later launch retries it
|
|
- `MaxConsecutiveTransient` consecutive transient failures aborts the pass rather than skipping
|
|
item-by-item against an edge that is refusing all traffic
|
|
- Only call `ApiKeyPool.MarkAvailable()` on a response that actually reached the API. A transport or
|
|
CDN failure says nothing about the key's standing and must not clear its flag
|
|
- Only a real HTTP 429 (or `meta.status == 429`) counts as a rate limit. Do not infer one from the
|
|
presence of `X-RateLimit-*` headers, which Tumblr sends on every response
|
|
- Rate limiters must pace with `await AcquireAsync()`. `AttemptAcquire()` does not wait, so a
|
|
saturated window aborts the run instead of throttling it
|
|
- Long-running commands return exit 3 when a pass ends incomplete (rate-limit pause, breaker trip, or
|
|
skipped items), so a caller can distinguish that from a clean run
|
|
|
|
### `Notes` Stores Integer IDs, Not Names
|
|
As of 2026-08-07 `Notes.RootBlogName`, `NoteBlogName` and `Type` are gone, replaced by
|
|
`RootBlogId`, `NoteBlogId` and `TypeId`. Blog IDs resolve through `Blogs.BlogId`, and
|
|
types through the `NoteTypes` lookup table. There is no compatibility view: naming an old
|
|
column is a hard SQLite error, so unlike `IsActive` this is a hard cut with no runtime
|
|
probe. Full detail in `URLNotesGrabberCORE/TL.db.md`.
|
|
|
|
- **`Blogs.BlogId` is the only blog-ID authority (since 2026-09-28).** IDs used to live in a
|
|
`BlogNames` table with an unmaintained copy in `Blogs.BlogId`. The copy drifted and hid
|
|
12k blogs from `GetBlogs`, so `retire-blognames.sql` moved the authority into `Blogs`
|
|
and **dropped `BlogNames` entirely**. There is no compatibility view, so naming it is
|
|
`no such table`. Do not recreate it
|
|
- **Joining `Notes` to `Blogs`**: `FROM Blogs B INNER JOIN Notes N ON N.NoteBlogId = B.BlogId`
|
|
- **Joining `Notes` to `Posts` also goes through `Blogs`**, since `Posts` has only
|
|
`BlogName`: `Posts P JOIN Blogs RB ON RB.BlogName = P.BlogName JOIN Notes N ON
|
|
N.RootBlogId = RB.BlogId` (`GetRepliesWithFilledText`)
|
|
- **Resolve a name by filtering `Blogs`, never by scanning `Notes`**:
|
|
`WHERE NoteBlogId = (SELECT BlogId FROM Blogs WHERE BlogName = @name)`. The subquery is a
|
|
primary-key probe and does not show against the 1.2M-row table
|
|
- **`AddNote` registers both blogs *and* the note type** before inserting, all in one
|
|
transaction. `RegisterBlog` does `INSERT OR IGNORE` into `Blogs`, then assigns
|
|
`BlogId = MAX(BlogId) + 1` where it is NULL. Unlike `AddBlog`, it does not skip `deact`
|
|
names, because a note by a deactivated blog still needs an ID. `NoteTypes` is a table
|
|
rather than a `CHECK` constraint precisely so an unseen type is an `INSERT`. Without
|
|
that registration a type would resolve to `NULL` and fail the `NOT NULL` on `TypeId`,
|
|
losing the note
|
|
- **Assigning a `BlogId` is bookkeeping and must not move `DateModified`**
|
|
- **`Blogs.BlogId` is NULL on ~166k of ~199k rows**, every blog that has never appeared in
|
|
a note. An inner join on it silently drops them. Correct for engagement queries, wrong
|
|
for anything listing the registry. `ix_Blogs_BlogId` is `UNIQUE`, which allows many NULLs
|
|
- **IDs are stable and must never be renumbered.** They are stored in 1.2M `Notes` rows.
|
|
Triggers `trg_Blogs_BlogId_NoDelete` and `trg_Blogs_BlogId_Immutable` abort any
|
|
`DELETE` of a `Blogs` row that has a `BlogId`, and any change to its `BlogId` or
|
|
`BlogName`. A blog renamed upstream gets a new row. Remove a blog with `IsActive = 0`.
|
|
These triggers are also what make `MAX(BlogId) + 1` safe: no ID can ever be freed for
|
|
reuse
|
|
- Prefer `TypeId = (SELECT TypeId FROM NoteTypes WHERE Type = 'reply')` over a hardcoded
|
|
ID. A negated `TypeId NOT IN (SELECT …)` is only correct because `TypeId` is `NOT NULL`
|
|
- Duplicate-key detection uses `IsNotesDuplicateKey`, which matches the constraint and the
|
|
table rather than an exact column list. The old literal string comparison broke silently
|
|
on this rename — do not reintroduce one
|
|
|
|
### `IsActive` Is Not Ours To Write
|
|
`Blogs.IsActive`, `Posts.IsActive` and `Notes.IsActive` are removal flags set by other tools
|
|
(Rolodex). `0` means removed; anything else, including `NULL`, means live. Full detail in
|
|
`URLNotesGrabberCORE/TL.db.md`.
|
|
|
|
- **Never write any `IsActive` column.** Not in an `INSERT` column list, not in an `UPDATE`,
|
|
and never via `INSERT OR REPLACE` on these tables — that resets the column default and
|
|
un-removes the row. Re-crawling a removed row must refresh its content and leave the flag
|
|
where it was
|
|
- **Filter at selection, not at write.** Every query that *selects* posts, notes or blogs
|
|
excludes removed rows. Update statements stay keyed on a row the caller already selected;
|
|
filtering them would spend API quota and then fail to persist the result
|
|
- `Posts.IsActive` and `Notes.IsActive` are **optional** — they are absent from databases
|
|
that predate them, and naming a missing column is a hard SQLite error. Compose the filter
|
|
with `AndIsActive`/`WhereIsActive` in `DataAccess.cs`, which return
|
|
`COALESCE(IsActive, 1) = 1` only when `HasIsActiveColumn` finds the column. `Blogs.IsActive`
|
|
is not optional and is filtered directly
|
|
- Do not add these columns from this app, and do not add them to the missing-column list in
|
|
`verify-db-schema.sql`
|
|
|
|
### `DateModified` Tracks Real Changes Only
|
|
`Blogs.DateModified`, `Posts.DateModified` and `Notes.DateModified` must move only when a
|
|
column beside `DateModified` itself actually changed. Re-crawling or re-ingesting identical
|
|
content has to leave the row — and its timestamp — untouched, or downstream consumers cannot
|
|
tell a refreshed row from a rewritten one.
|
|
|
|
- Enforce it in the `WHERE` clause, not in C#. Every `UPDATE` that sets `DateModified` ends
|
|
with an `AND (<col> <> @param OR ...)` term covering every column in its `SET` list, so
|
|
SQLite matches zero rows on a no-op and never writes
|
|
- Compare NULL-safely: `IFNULL(col, '') <> IFNULL(@param, '')` for text,
|
|
`IFNULL(col, 0) <> @param` for integer flags. A bare `col <> @param` is NULL on a NULL
|
|
column and silently skips the row that most needs writing
|
|
- Where NULL is not equivalent to the default, spell it out. The `HasBeenOutput = 0` stamps
|
|
use `(HasBeenOutput IS NULL OR HasBeenOutput <> 0)` because the selection queries test
|
|
`HasBeenOutput = 0`, which a NULL would never match
|
|
- Dynamic `SET` lists (`UpdatePostContentFields`) build the guard alongside the assignments
|
|
so the two lists cannot drift apart
|
|
- These statements now return 0 rows for "found but unchanged" as well as "not found".
|
|
Callers that read `ExecuteNonQuery()` must not treat 0 as "row missing"
|
|
|
|
**`Posts.NotesGatheredDateTime` is crawl bookkeeping, not content.** It moves on every
|
|
`-collect` pass and says nothing about the post, so it must never move `DateModified` on its
|
|
own. `UpdatePostMarkNotesCollected` still writes it every pass but wraps the timestamp in
|
|
`DateModified = CASE WHEN IFNULL(HasNotesGathered, 0) <> 1 THEN @dateModified ELSE
|
|
DateModified END` — SQLite evaluates `SET` expressions against the pre-`UPDATE` row, so only
|
|
the flag flipping counts as a modification. Use this shape for any column that has to be
|
|
refreshed unconditionally without being a change. `Blogs.LikesLastRefreshed` is the
|
|
deliberate exception: a refresh pass is treated as a real event on the blog row.
|
|
|
|
**`Blogs.DateAdded` is write-once.** `AddBlog`'s `INSERT` is the only place that sets it. A
|
|
new post arriving for a known blog reopens `HasBeenOutput` but must leave `DateAdded` alone —
|
|
a new post is not a new blog, and rewriting the column both destroys the registration date
|
|
and makes every insert look like a change.
|
|
|
|
**`"."` in a `Posts` content field means "not supplied", not "empty".** `ReblogRecord`
|
|
(`TraverseDirectory`'s parser for the local `.txt` export tree) and the `--likes` API path
|
|
both default every content field to the literal string `"."` when their source has no value
|
|
for it, then pass that straight to `UpdatePost`. A blog with two export folders in different
|
|
field formats (a duplicate `_2` folder, or a Tumblr export whose field set changed over time)
|
|
sends one record with a real `Title`/`Tags`/`Slug` and another with those fields `"."`
|
|
because that format never had a line for them — and without a guard, re-importing both on
|
|
every run flips the row back and forth forever, bumping `DateModified` on every pass even
|
|
though the true content never changes.
|
|
|
|
- Every content column in `UpdatePost`'s `SET` list is guarded the same way `RootBlogName`/
|
|
`RootURL` already were: `col = CASE WHEN @col = '.' THEN col ELSE @col END`. A `"."`
|
|
parameter leaves the existing value alone instead of overwriting it
|
|
- The change-detection `WHERE` clause carries the same exception —
|
|
`(@col <> '.' AND IFNULL(col, '') <> @col) OR ...` — so a `"."`-only difference does not
|
|
make the statement fire at all, and `DateModified` stays put
|
|
- Deliberately narrow: only the literal `"."` is the sentinel. An explicit empty string from
|
|
a real record still overwrites, same as before this fix. `postID`, `BlogName`, `hasImage`,
|
|
`ByLikes` are not part of this convention and are unaffected
|
|
- If a new content field is added to `Posts`/`UpdatePost`, decide explicitly whether its
|
|
source can legitimately supply `"."` as "field absent" before deciding whether it needs
|
|
the same `CASE` treatment — don't assume every column needs it
|
|
|
|
**`--ingest` (`UpsertPostFromTextFile`) uses `NULL`, not `"."`, for the same "field absent"
|
|
convention, and reconciling exactly this kind of duplicate IS the feature's job.**
|
|
`IngestMode` strips a trailing `_N` from the folder name before it ever reaches
|
|
`UpsertPostFromTextFile`, so a duplicate export folder collapses onto the same `BlogName` on
|
|
purpose — the whole point is to merge multiple differently-formatted files for the same post
|
|
into one row. `IngestMode.G(key)` returns `null` (not `"."`) when a field's line is absent
|
|
from a given file, `LegacyPostsDbImporter` passes `null` straight from a `NULL` source column,
|
|
and files are walked in raw filesystem enumeration order — never sorted — so which file's call
|
|
lands last for a given `(BlogName, PostID)` is arbitrary.
|
|
|
|
- Before the fix, the `UPDATE` branch set every column unconditionally, so whichever file
|
|
processed last for a `PostID` would null out every field its own record didn't carry —
|
|
silently erasing real `Title`/`Slug`/`Tags`/… another file had, the opposite of what
|
|
`--ingest` exists to do. This is worse than the `"."` case above: that one only caused
|
|
churn (the two writes canceled out); this one loses data, and which posts lose which
|
|
fields depends on filesystem enumeration order
|
|
- Same shape of fix, `NULL` instead of `"."` as the sentinel: `col = CASE WHEN @col IS NULL
|
|
THEN col ELSE @col END` in the `SET` list, `(@col IS NOT NULL AND IFNULL(col, '') <> @col)
|
|
OR ...` in the change-detection
|
|
- Same narrow rule: only `NULL` (the field's line was never present in this file) is the
|
|
sentinel. `G()` already distinguishes this from "present but blank" — a dictionary miss is
|
|
`null`, an empty value after the prefix is `""` — so an explicitly blank field still
|
|
overwrites
|
|
- `HasImage` is **not** guarded and remains a known gap: `IngestMode` always computes a
|
|
concrete `bool` (defaulting `false` when a file has no `Has Image:` line), so there is no
|
|
way for this function to tell "this format says no image" from "this format doesn't report
|
|
it at all" without changing the parameter to `bool?` and threading that through
|
|
`IngestMode`/`LegacyPostsDbImporter`. Fix this the same way if `--ingest` is observed
|
|
downgrading a post's `HasImage` from `1` to `0`
|
|
|
|
### Testing
|
|
- No existing test suite; use xUnit if adding tests
|
|
- Test critical logic: `ApiKeyPool` init, color parsing, config persistence
|
|
- Avoid testing one-off CLI workflows
|
|
|
|
### Git Commit Messages
|
|
- Imperative mood ("Add feature" not "Added feature")
|
|
- Prefix with type: `feat:`, `fix:`, `chore:`, `docs:`
|
|
- Keep messages under 72 characters
|
|
- Never commit sensitive data (API keys/tokens)
|
|
|
|
### API Key Color Rules
|
|
- Unconfigured keys auto-assign colors from a preset palette
|
|
- Auto-assigned colors persist to `appsettings.json`
|
|
- All output for an active key uses its assigned color
|
|
- Temporary color changes (e.g., errors) must restore the key's color afterward
|