fix(db): make Blogs.BlogId the only blog ID and drop BlogNames

Blogs.BlogId was a one-time copy of BlogNames and nothing kept it
current: 12,238 blogs first seen after 2026-08-07 had a BlogNames ID
but a NULL Blogs.BlogId, so GetBlogs' join on BlogId silently skipped
them and their 23,148 notes.

- retire-blognames.sql: stub Blogs rows for the 17 unregistered note
  participants, backfill IDs (none renumbered), make ix_Blogs_BlogId
  UNIQUE, drop BlogNames, and add triggers that stop a Blogs row with
  a BlogId from being deleted, renamed or renumbered
- AddNote registers both blogs via RegisterBlog (Blogs row + MAX+1 ID)
  and every query resolves names through Blogs instead of BlogNames
- verify-db-schema.sql reports a DB that still has BlogNames (1e)
- Update TL.db.md, AGENTS.md and the DB Browser saved queries

Co-Authored-By: Claude Opus 5.5 <[email protected]>
This commit is contained in:
jim
2026-09-28 11:33:04 -05:00
co-authored by Claude Opus 5.5
parent c8c43c4918
commit 8fe2ffeb96
8 changed files with 751 additions and 232 deletions
+33 -21
View File
@@ -49,28 +49,40 @@ say nothing about the item being fetched, so they must not be recorded as per-it
### `Notes` Stores Integer IDs, Not Names
As of 2026-08-07 `Notes.RootBlogName`, `NoteBlogName` and `Type` are gone, replaced by
`RootBlogId`, `NoteBlogId` and `TypeId` resolving through the `BlogNames` and `NoteTypes`
lookup tables. There is no compatibility view — naming an old column is a hard SQLite
error, so unlike `IsActive` this is a hard cut with no runtime probe. Full detail in
`URLNotesGrabberCORE/TL.db.md`.
`RootBlogId`, `NoteBlogId` and `TypeId`. Blog IDs resolve through `Blogs.BlogId`, and
types through the `NoteTypes` lookup table. There is no compatibility view: naming an old
column is a hard SQLite error, so unlike `IsActive` this is a hard cut with no runtime
probe. Full detail in `URLNotesGrabberCORE/TL.db.md`.
- **Joining `Notes` to `Blogs` goes through `Blogs.BlogId`**, not `BlogNames`:
`FROM Blogs B INNER JOIN Notes N ON N.NoteBlogId = B.BlogId`. Routing it through
`BlogNames` adds a hop and ends in the text comparison the migration removed
- **Joining `Notes` to `Posts` is the opposite** — `Posts` has only `BlogName`, so it must
go through `BlogNames` (`GetRepliesWithFilledText`). This is the only such join
- **Resolve a name by filtering the lookup, never by scanning `Notes`**:
`WHERE NoteBlogId = (SELECT BlogId FROM BlogNames WHERE BlogName = @name)`. The subquery
is a unique-index probe on 20k rows and does not show against the 1.18M-row table
- **`AddNote` registers both blog names *and* the note type** with `INSERT OR IGNORE`
before inserting, all in one transaction. `NoteTypes` is a table rather than a `CHECK`
constraint precisely so an unseen type is an `INSERT`; without that registration it
would resolve to `NULL` and fail the `NOT NULL` on `TypeId`, losing the note
- **`Blogs.BlogId` is NULL on 168,202 of 188,620 rows** — every blog that has never
appeared in a note. An inner join on it silently drops them. Correct for engagement
queries, wrong for anything listing the registry
- **IDs are stable and must never be renumbered.** They are stored in 1.18M `Notes` rows.
A blog renamed upstream gets a new `BlogNames` row, not an edited one
- **`Blogs.BlogId` is the only blog-ID authority (since 2026-09-28).** IDs used to live in a
`BlogNames` table with an unmaintained copy in `Blogs.BlogId`. The copy drifted and hid
12k blogs from `GetBlogs`, so `retire-blognames.sql` moved the authority into `Blogs`
and **dropped `BlogNames` entirely**. There is no compatibility view, so naming it is
`no such table`. Do not recreate it
- **Joining `Notes` to `Blogs`**: `FROM Blogs B INNER JOIN Notes N ON N.NoteBlogId = B.BlogId`
- **Joining `Notes` to `Posts` also goes through `Blogs`**, since `Posts` has only
`BlogName`: `Posts P JOIN Blogs RB ON RB.BlogName = P.BlogName JOIN Notes N ON
N.RootBlogId = RB.BlogId` (`GetRepliesWithFilledText`)
- **Resolve a name by filtering `Blogs`, never by scanning `Notes`**:
`WHERE NoteBlogId = (SELECT BlogId FROM Blogs WHERE BlogName = @name)`. The subquery is a
primary-key probe and does not show against the 1.2M-row table
- **`AddNote` registers both blogs *and* the note type** before inserting, all in one
transaction. `RegisterBlog` does `INSERT OR IGNORE` into `Blogs`, then assigns
`BlogId = MAX(BlogId) + 1` where it is NULL. Unlike `AddBlog`, it does not skip `deact`
names, because a note by a deactivated blog still needs an ID. `NoteTypes` is a table
rather than a `CHECK` constraint precisely so an unseen type is an `INSERT`. Without
that registration a type would resolve to `NULL` and fail the `NOT NULL` on `TypeId`,
losing the note
- **Assigning a `BlogId` is bookkeeping and must not move `DateModified`**
- **`Blogs.BlogId` is NULL on ~166k of ~199k rows**, every blog that has never appeared in
a note. An inner join on it silently drops them. Correct for engagement queries, wrong
for anything listing the registry. `ix_Blogs_BlogId` is `UNIQUE`, which allows many NULLs
- **IDs are stable and must never be renumbered.** They are stored in 1.2M `Notes` rows.
Triggers `trg_Blogs_BlogId_NoDelete` and `trg_Blogs_BlogId_Immutable` abort any
`DELETE` of a `Blogs` row that has a `BlogId`, and any change to its `BlogId` or
`BlogName`. A blog renamed upstream gets a new row. Remove a blog with `IsActive = 0`.
These triggers are also what make `MAX(BlogId) + 1` safe: no ID can ever be freed for
reuse
- Prefer `TypeId = (SELECT TypeId FROM NoteTypes WHERE Type = 'reply')` over a hardcoded
ID. A negated `TypeId NOT IN (SELECT …)` is only correct because `TypeId` is `NOT NULL`
- Duplicate-key detection uses `IsNotesDuplicateKey`, which matches the constraint and the