diff --git a/AGENTS.md b/AGENTS.md
index 066e933..04445fc 100644
--- a/AGENTS.md
+++ b/AGENTS.md
@@ -49,28 +49,40 @@ say nothing about the item being fetched, so they must not be recorded as per-it
### `Notes` Stores Integer IDs, Not Names
As of 2026-08-07 `Notes.RootBlogName`, `NoteBlogName` and `Type` are gone, replaced by
-`RootBlogId`, `NoteBlogId` and `TypeId` resolving through the `BlogNames` and `NoteTypes`
-lookup tables. There is no compatibility view — naming an old column is a hard SQLite
-error, so unlike `IsActive` this is a hard cut with no runtime probe. Full detail in
-`URLNotesGrabberCORE/TL.db.md`.
+`RootBlogId`, `NoteBlogId` and `TypeId`. Blog IDs resolve through `Blogs.BlogId`, and
+types through the `NoteTypes` lookup table. There is no compatibility view: naming an old
+column is a hard SQLite error, so unlike `IsActive` this is a hard cut with no runtime
+probe. Full detail in `URLNotesGrabberCORE/TL.db.md`.
-- **Joining `Notes` to `Blogs` goes through `Blogs.BlogId`**, not `BlogNames`:
- `FROM Blogs B INNER JOIN Notes N ON N.NoteBlogId = B.BlogId`. Routing it through
- `BlogNames` adds a hop and ends in the text comparison the migration removed
-- **Joining `Notes` to `Posts` is the opposite** — `Posts` has only `BlogName`, so it must
- go through `BlogNames` (`GetRepliesWithFilledText`). This is the only such join
-- **Resolve a name by filtering the lookup, never by scanning `Notes`**:
- `WHERE NoteBlogId = (SELECT BlogId FROM BlogNames WHERE BlogName = @name)`. The subquery
- is a unique-index probe on 20k rows and does not show against the 1.18M-row table
-- **`AddNote` registers both blog names *and* the note type** with `INSERT OR IGNORE`
- before inserting, all in one transaction. `NoteTypes` is a table rather than a `CHECK`
- constraint precisely so an unseen type is an `INSERT`; without that registration it
- would resolve to `NULL` and fail the `NOT NULL` on `TypeId`, losing the note
-- **`Blogs.BlogId` is NULL on 168,202 of 188,620 rows** — every blog that has never
- appeared in a note. An inner join on it silently drops them. Correct for engagement
- queries, wrong for anything listing the registry
-- **IDs are stable and must never be renumbered.** They are stored in 1.18M `Notes` rows.
- A blog renamed upstream gets a new `BlogNames` row, not an edited one
+- **`Blogs.BlogId` is the only blog-ID authority (since 2026-09-28).** IDs used to live in a
+ `BlogNames` table with an unmaintained copy in `Blogs.BlogId`. The copy drifted and hid
+ 12k blogs from `GetBlogs`, so `retire-blognames.sql` moved the authority into `Blogs`
+ and **dropped `BlogNames` entirely**. There is no compatibility view, so naming it is
+ `no such table`. Do not recreate it
+- **Joining `Notes` to `Blogs`**: `FROM Blogs B INNER JOIN Notes N ON N.NoteBlogId = B.BlogId`
+- **Joining `Notes` to `Posts` also goes through `Blogs`**, since `Posts` has only
+ `BlogName`: `Posts P JOIN Blogs RB ON RB.BlogName = P.BlogName JOIN Notes N ON
+ N.RootBlogId = RB.BlogId` (`GetRepliesWithFilledText`)
+- **Resolve a name by filtering `Blogs`, never by scanning `Notes`**:
+ `WHERE NoteBlogId = (SELECT BlogId FROM Blogs WHERE BlogName = @name)`. The subquery is a
+ primary-key probe and does not show against the 1.2M-row table
+- **`AddNote` registers both blogs *and* the note type** before inserting, all in one
+ transaction. `RegisterBlog` does `INSERT OR IGNORE` into `Blogs`, then assigns
+ `BlogId = MAX(BlogId) + 1` where it is NULL. Unlike `AddBlog`, it does not skip `deact`
+ names, because a note by a deactivated blog still needs an ID. `NoteTypes` is a table
+ rather than a `CHECK` constraint precisely so an unseen type is an `INSERT`. Without
+ that registration a type would resolve to `NULL` and fail the `NOT NULL` on `TypeId`,
+ losing the note
+- **Assigning a `BlogId` is bookkeeping and must not move `DateModified`**
+- **`Blogs.BlogId` is NULL on ~166k of ~199k rows**, every blog that has never appeared in
+ a note. An inner join on it silently drops them. Correct for engagement queries, wrong
+ for anything listing the registry. `ix_Blogs_BlogId` is `UNIQUE`, which allows many NULLs
+- **IDs are stable and must never be renumbered.** They are stored in 1.2M `Notes` rows.
+ Triggers `trg_Blogs_BlogId_NoDelete` and `trg_Blogs_BlogId_Immutable` abort any
+ `DELETE` of a `Blogs` row that has a `BlogId`, and any change to its `BlogId` or
+ `BlogName`. A blog renamed upstream gets a new row. Remove a blog with `IsActive = 0`.
+ These triggers are also what make `MAX(BlogId) + 1` safe: no ID can ever be freed for
+ reuse
- Prefer `TypeId = (SELECT TypeId FROM NoteTypes WHERE Type = 'reply')` over a hardcoded
ID. A negated `TypeId NOT IN (SELECT …)` is only correct because `TypeId` is `NOT NULL`
- Duplicate-key detection uses `IsNotesDuplicateKey`, which matches the constraint and the
diff --git a/RERUN.sqbpro b/RERUN.sqbpro
index 0134308..a05f74e 100644
--- a/RERUN.sqbpro
+++ b/RERUN.sqbpro
@@ -1,59 +1,69 @@
-UPDATE Posts
-SET HasNotesGathered = 0
-WHERE (BlogName, PostID) IN (
- SELECT p.BlogName, p.PostID
- FROM Posts p
- WHERE p.HasNotesGathered = 1
- AND P.notesGatheredDatetime < 1774294520
- AND EXISTS (
- SELECT 1
- FROM Notes n
- WHERE n.PostID = p.PostID
- AND n.RootBlogId = (SELECT BlogId FROM BlogNames WHERE BlogName = p.BlogName)
- --AND n.TypeId NOT IN (SELECT TypeId FROM NoteTypes WHERE Type IN ('reblog', 'reply'))
- )
- ORDER BY P.PostDate ASC
- --LIMIT 500
-);select *
+select *
from Blogs
--update blogs set HasBeenOutput = 1
where HasBeenOutput = 0
AND
blogname in
-(
-'teaberrybee',
-'reddevilgoddesstoo',
-'waywardog13',
-'wzjustbrowsing-blog',
-'lewerta',
-'nudenymph',
-'caylachief'
-
+('udontn33dh1m',
+'tyrantsxblood',
+'sentry-34',
+'deathcabforfrankie',
+'abheith-sasta',
+'kuwaiikittenghost',
+'kansasmud',
+'03diesel',
+'itzameallieee',
+'fireball-temptations',
+'mamaisamess',
+'906raised-and-dogobsessed',
+'the-queerist-wolf',
+'counting-corpsess',
+'aqueenbby',
+'maybememoriesx',
+'queenofnevers',
+'obsidian-psyche',
+'lilmissellexo',
+'alittlebunny95',
+'rage--and--grace',
+'savage-deniz',
+'daddyspuddleprincess',
+'littledefenstration',
+'bearded-snorlax',
+'thosesummerskiess',
+'tubadtoph',
+'lieutenant-dan-ice-cream',
+'brittvnybitch',
+'a-smol-gayologist',
+'sum1random',
+'samsternelly',
+'littlemouseylauren',
+'princessleiaorgasma',
+'bloodstaineddkisses',
+'letsfacerealitybabe',
+'x--marks--thespot',
+'space-and-suffering',
+'rinarootski',
+'thiccandtired',
+'fvcking-scvmbag',
+'fullblownwizard',
+'bigjewface',
+'unleash-the-krayken',
+'bumpintheroad',
+'liltexasjedii',
+'nawtydude',
+'queenpeachqueen',
+'the-clansman',
+'balmain-bxtch'
)select P.slug, N.replyText, rbn.BlogName as RootBlogName, n.PostID, nbn.BlogName || '.tumblr.com' as NoteBlogName, DatetimeCrawled, TimeStamp, nt.Type, rbn.BlogName || '.tumblr.com/post/' || n.postid, datetime(timestamp, 'unixepoch')
from Notes N
inner join Posts P on p.PostID = n.PostID
-inner join BlogNames rbn on rbn.BlogId = n.RootBlogId
-inner join BlogNames nbn on nbn.BlogId = n.NoteBlogId
+inner join Blogs rbn on rbn.BlogId = n.RootBlogId
+inner join Blogs nbn on nbn.BlogId = n.NoteBlogId
inner join NoteTypes nt on nt.TypeId = n.TypeId
where
DatetimeCrawled > '2026-08-07 11:47:22' and nt.Type like 'r%'
and P.IsActive = 1
-order by n.DatetimeCrawledSELECT distinct
- '''' || blogname || ''',',
- blogs.*
- , blogname || '.tumblr.com'
-FROM
- Blogs
- inner JOIN
- Notes on notes.noteBlogId = blogs.BlogId
- inner JOIN
- NoteTypes on NoteTypes.TypeId = Notes.TypeId
-WHERE
- HasBeenOutput = 0 and NoteTypes.Type = 'reblog'
-order by
- NoteTypes.Type desc,
- DateAdded desc
-LIMIT 100;WITH ReplyCounts AS (
+order by n.DatetimeCrawledWITH ReplyCounts AS (
SELECT
NoteBlogId,
COUNT(DISTINCT replyText) AS DistinctReplyCount
@@ -68,13 +78,13 @@ SELECT
c.DistinctReplyCount
FROM Notes n
JOIN ReplyCounts c ON n.NoteBlogId = c.NoteBlogId
-JOIN BlogNames rbn ON rbn.BlogId = n.RootBlogId
-JOIN BlogNames nbn ON nbn.BlogId = n.NoteBlogId
+JOIN Blogs rbn ON rbn.BlogId = n.RootBlogId
+JOIN Blogs nbn ON nbn.BlogId = n.NoteBlogId
JOIN NoteTypes t ON t.TypeId = n.TypeId
where replyText <> '.' and t.Type <> 'reply'
--AND nbn.BlogName NOT IN ( 'roadblocker21', 'thesaddemon666', 'edwardabbeyhoffman', 'tattedsoldier20', 'zomb-eh', 'animalistic13', 'indken', 'maccloud1592',
--'moss-wizard', 'supertrucker12682', 'exploringthrupics', 'padeyepete' )
-order by c.DistinctReplyCount desc, nbn.BlogName, n.DateModified desc, replyText, rbn.BlogName, PostIDWITH PostsWithCount AS ( SELECT P.BlogName, P.PostID, 1925013599 AS LatestNoteTimestamp, P.NotesGatheredDateTime, COUNT(P.PostID) OVER(PARTITION BY P.BlogName) AS CNT, P.HasNotesGathered, P.NotFound, P.PostDate FROM Posts P WHERE COALESCE(P.IsActive, 1) = 1 ), Unioned AS ( SELECT BlogName, PostID, LatestNoteTimestamp, NotesGatheredDateTime, CNT, PostDate FROM PostsWithCount WHERE NotFound = 0 AND HasNotesGathered = 0 UNION SELECT BlogName, PostID, LatestNoteTimestamp, NotesGatheredDateTime, CNT, PostDate FROM PostsWithCount WHERE BlogName = 'zomb-eh' AND NotFound = 0 AND NotesGatheredDateTime < unixepoch('now', 'localtime', '-3 days') ) SELECT U.BlogName, U.PostID, U.LatestNoteTimestamp, U.NotesGatheredDateTime, U.CNT FROM Unioned U WHERE (U.NotesGatheredDateTime < 1786134037 OR U.NotesGatheredDateTime IS NULL) ORDER BY U.NotesGatheredDateTime, U.PostDate DESC, U.BlogName, U.PostID;delete from posts where postid in
+order by c.DistinctReplyCount desc, nbn.BlogName, n.DateModified desc, replyText, rbn.BlogName, PostIDWITH PostsWithCount AS ( SELECT P.BlogName, P.PostID, 1925013599 AS LatestNoteTimestamp, P.NotesGatheredDateTime, COUNT(P.PostID) OVER(PARTITION BY P.BlogName) AS CNT, P.HasNotesGathered, P.NotFound, P.PostDate FROM Posts P WHERE COALESCE(P.IsActive, 1) = 1 ), Unioned AS ( SELECT BlogName, PostID, LatestNoteTimestamp, NotesGatheredDateTime, CNT, PostDate FROM PostsWithCount WHERE NotFound = 0 AND HasNotesGathered = 0 UNION SELECT BlogName, PostID, LatestNoteTimestamp, NotesGatheredDateTime, CNT, PostDate FROM PostsWithCount WHERE BlogName = 'zomb-eh' AND NotFound = 0 AND NotesGatheredDateTime < unixepoch('now', 'localtime', '-3 days') ) SELECT U.BlogName, U.PostID, U.LatestNoteTimestamp, U.NotesGatheredDateTime, U.CNT FROM Unioned U WHERE (U.NotesGatheredDateTime < 1787237598 OR U.NotesGatheredDateTime IS NULL) ORDER BY U.NotesGatheredDateTime, U.PostDate DESC, U.BlogName, U.PostID;delete from posts where postid in
(
'741662499571728384',
178892849664,
@@ -82,28 +92,324 @@ order by c.DistinctReplyCount desc, nbn.BlogName, n.DateModified desc, replyText
177012868749,
169950081964,
755440787056099328
-)select *
+)select *
-- delete
from notes
-where postid not in (select distinct postid from posts where IsActive = 1)SELECT
- *
-FROM
- POSTS P
-WHERE
- P.ByLikes = 1
- AND
- P.DateCreated > '2026-05-26 17:47:32'
-ORDER BY
- P.DateCreated descupdate posts set IsActive = 0 where blogname IN ( 'shoebiedoo', 'redheaded-girlygirl', 'xlittle-ghost' )update Posts
-set IsActive = 0
-where postid in
-(
-
-
-'731937314675310592'
-
+where postid not in (select distinct postid from posts where IsActive = 1)-- ============================================================================
+-- blogs-added-after-august-2026-with-reblog-or-reply.sql
+--
+-- Purpose: Of all blogs in Blogs, find the ones that show up in Notes as the
+-- engager (NoteBlogId) on a 'reblog' or 'reply' note, sorted by
+-- DateAdded. (Originally scoped to "added after August 2026" --
+-- that cutoff is now removed per request; QUERY 2 shows how to put
+-- a date floor back if needed.)
+--
+-- Read-only. No INSERT/UPDATE/DELETE/DDL anywhere in this file.
+--
+-- How to use (DB Browser for SQLite):
+-- 1. File > Open Database -> TL.db
+-- 2. Execute SQL tab. Each QUERY below is independent; Ctrl+Enter runs just
+-- the one your cursor is in.
+--
+-- The join, once:
+-- "Added after August 2026" filters Blogs.DateAdded. "Has a reblog/reply
+-- note" means the blog is the engager, which is NoteBlogId -- not
+-- RootBlogId, which is the blog that *owns* the post being reacted to
+-- (see find-notes-on-inactive-posts.sql for that side). Per TL.db.md, the
+-- Blogs<->Notes join is a single integer hop and should not be routed
+-- through Blogs:
+-- Blogs.BlogId = Notes.NoteBlogId
+-- EXISTS is used rather than a JOIN so a blog with many qualifying notes
+-- still contributes one output row.
+--
+-- Excluding notes on an inactive post: same shape as
+-- find-notes-on-inactive-posts.sql -- Notes only carries RootBlogId (an
+-- integer), so reaching Posts.IsActive needs the one text hop the rest of
+-- this file avoids: RootBlogId -> Blogs.BlogName = Posts.BlogName,
+-- matched on PostID. It's a LEFT JOIN, not an inner one: only 3,867 of
+-- 20,430 root/engager blogs have any stored Posts rows at all (TL.db.md),
+-- so most reblog/reply notes have no Posts row to check and must be kept,
+-- not dropped by an inner join. COALESCE(p.IsActive, 1) = 1 keeps a note
+-- unless its post is explicitly IsActive = 0 -- NULL (no Posts row, or a
+-- stored row with no flag written) means live, per the schema's own
+-- convention (TL.db.md, "Posts.IsActive and Notes.IsActive"). This is a
+-- big filter in practice: of the blogs that qualified before it, most
+-- have every one of their reblog/reply notes pointing at a since-removed
+-- post, not just some -- verified against the live data, not assumed.
+--
+-- On DateAdded: this column is not written consistently -- most rows hold
+-- ISO 'yyyy-MM-dd HH:mm:ss', but 17k+ hold US 'M/d/yy' from a 2025 bulk
+-- import (see TL.db.md, "DateAdded is not written consistently"). As text,
+-- those two shapes do not sort or compare against each other correctly, so
+-- QUERY 0 normalises both to an ISO date before filtering. In the live data
+-- every US-format row predates August 2026 anyway (only '12/23/25' and
+-- '12/24/25' occur), so this makes no difference to the current answer --
+-- it's here so the query stays correct if that ever changes.
+-- ============================================================================
+-- ----------------------------------------------------------------------------
+-- QUERY 0 / THE ANSWER -- one row per qualifying blog, sorted by DateAdded
+-- descending (normalised -- see the note above). No date cutoff, but now
+-- scoped to HasBeenOutput = 0 AND IsActive = 1. 4,739 rows in the live
+-- data.
+--
+-- earliest_reblog_or_reply_utc is the MIN(TimeStamp) among this blog's
+-- reblog-or-reply notes (either type counts -- see the column name).
+-- Getting this meant switching QUERY 0 from EXISTS to an inner JOIN +
+-- GROUP BY: EXISTS can only tell you a qualifying row is present, not
+-- aggregate over which ones. No CASE is needed inside the MIN() because
+-- the WHERE below already restricts the joined rows to reblog/reply, so
+-- every row a blog brings into the aggregate is one this column should
+-- consider. A blog appears exactly once, same as before, and this column
+-- is never NULL for a row that's in the result at all (an earlier
+-- revision aggregated reblog only, which left it NULL for the 181 blogs
+-- that had replies but no reblogs).
+-- ----------------------------------------------------------------------------
+WITH BlogsSplit AS (
+ SELECT
+ b.BlogId,
+ b.BlogName,
+ b.DateAdded,
+ CASE WHEN b.DateAdded LIKE '____-__-__%' THEN 1 ELSE 0 END AS IsIso,
+ -- for the US 'M/d/yy' shape only: everything after the first '/'
+ substr(b.DateAdded, instr(b.DateAdded, '/') + 1) AS RestAfterMonth
+ FROM Blogs b
+ WHERE b.BlogId IS NOT NULL and HasBeenOutput = 0 and IsActive = 1 -- a blog can only match Notes if it has one
+),
+BlogsNorm AS (
+ SELECT
+ BlogId,
+ BlogName,
+ DateAdded,
+ CASE
+ WHEN IsIso = 1 THEN date(DateAdded)
+ ELSE date(
+ '20' || substr(RestAfterMonth, instr(RestAfterMonth, '/') + 1) || '-' ||
+ substr('00' || substr(DateAdded, 1, instr(DateAdded, '/') - 1), -2) || '-' ||
+ substr('00' || substr(RestAfterMonth, 1, instr(RestAfterMonth, '/') - 1), -2)
+ )
+ END AS DateAddedNorm
+ FROM BlogsSplit
)
+SELECT
+ bn.BlogId,
+ bn.BlogName,
+ bn.DateAdded,
+ bn.DateAddedNorm,
+ datetime(MIN(n.TimeStamp), 'unixepoch') AS earliest_reblog_or_reply_utc
+ FROM BlogsNorm bn
+ JOIN Notes n ON n.NoteBlogId = bn.BlogId
+ JOIN NoteTypes t ON t.TypeId = n.TypeId
+ JOIN Blogs root_bn ON root_bn.BlogId = n.RootBlogId
+ LEFT JOIN Posts p ON p.BlogName = root_bn.BlogName
+ AND p.PostID = n.PostID
+ WHERE t.Type IN ('reblog')--, 'reply')
+ AND COALESCE(p.IsActive, 1) = 1 -- exclude notes on a post explicitly marked removed
+ GROUP BY bn.BlogId, bn.BlogName, bn.DateAdded, bn.DateAddedNorm
+ ORDER BY bn.DateAddedNorm desc;
-
+
+-- ----------------------------------------------------------------------------
+-- QUERY 1 -- same answer, with a per-blog breakdown of which type(s) fired
+-- and how many. Useful once QUERY 0 has rows; redundant while it's empty.
+-- ----------------------------------------------------------------------------
+-- WITH BlogsSplit AS ( ... ), BlogsNorm AS ( ... ) -- reuse the CTEs above
+--
+-- SELECT
+-- bn.BlogId,
+-- bn.BlogName,
+-- bn.DateAddedNorm,
+-- SUM(CASE WHEN t.Type = 'reblog' THEN 1 ELSE 0 END) AS reblog_count,
+-- SUM(CASE WHEN t.Type = 'reply' THEN 1 ELSE 0 END) AS reply_count
+-- FROM BlogsNorm bn
+-- JOIN Notes n ON n.NoteBlogId = bn.BlogId
+-- JOIN NoteTypes t ON t.TypeId = n.TypeId
+-- WHERE t.Type IN ('reblog', 'reply')
+-- GROUP BY bn.BlogId, bn.BlogName, bn.DateAddedNorm
+-- ORDER BY bn.DateAddedNorm;
+
+
+-- ----------------------------------------------------------------------------
+-- QUERY 2 -- put a date floor back, if wanted later.
+-- Same as QUERY 0, with one extra line in the outer WHERE:
+-- AND bn.DateAddedNorm >= '2026-09-01' -- or whatever cutoff
+-- ----------------------------------------------------------------------------
+-- ============================================================================
+-- blogs-added-after-august-2026-with-reblog-or-reply.sql
+--
+-- Purpose: Of all blogs in Blogs, find the ones that show up in Notes as the
+-- engager (NoteBlogId) on a 'reblog' or 'reply' note, sorted by
+-- DateAdded. (Originally scoped to "added after August 2026" --
+-- that cutoff is now removed per request; QUERY 2 shows how to put
+-- a date floor back if needed.)
+--
+-- Read-only. No INSERT/UPDATE/DELETE/DDL anywhere in this file.
+--
+-- How to use (DB Browser for SQLite):
+-- 1. File > Open Database -> TL.db
+-- 2. Execute SQL tab. Each QUERY below is independent; Ctrl+Enter runs just
+-- the one your cursor is in.
+--
+-- The join, once:
+-- "Added after August 2026" filters Blogs.DateAdded. "Has a reblog/reply
+-- note" means the blog is the engager, which is NoteBlogId -- not
+-- RootBlogId, which is the blog that *owns* the post being reacted to
+-- (see find-notes-on-inactive-posts.sql for that side). Per TL.db.md, the
+-- Blogs<->Notes join is a single integer hop and should not be routed
+-- through Blogs:
+-- Blogs.BlogId = Notes.NoteBlogId
+-- EXISTS is used rather than a JOIN so a blog with many qualifying notes
+-- still contributes one output row.
+--
+-- Excluding notes on an inactive post: same shape as
+-- find-notes-on-inactive-posts.sql -- Notes only carries RootBlogId (an
+-- integer), so reaching Posts.IsActive needs the one text hop the rest of
+-- this file avoids: RootBlogId -> Blogs.BlogName = Posts.BlogName,
+-- matched on PostID. It's a LEFT JOIN, not an inner one: only 3,867 of
+-- 20,430 root/engager blogs have any stored Posts rows at all (TL.db.md),
+-- so most reblog/reply notes have no Posts row to check and must be kept,
+-- not dropped by an inner join. COALESCE(p.IsActive, 1) = 1 keeps a note
+-- unless its post is explicitly IsActive = 0 -- NULL (no Posts row, or a
+-- stored row with no flag written) means live, per the schema's own
+-- convention (TL.db.md, "Posts.IsActive and Notes.IsActive"). This is a
+-- big filter in practice: of the blogs that qualified before it, most
+-- have every one of their reblog/reply notes pointing at a since-removed
+-- post, not just some -- verified against the live data, not assumed.
+--
+-- On DateAdded: this column is not written consistently -- most rows hold
+-- ISO 'yyyy-MM-dd HH:mm:ss', but 17k+ hold US 'M/d/yy' from a 2025 bulk
+-- import (see TL.db.md, "DateAdded is not written consistently"). As text,
+-- those two shapes do not sort or compare against each other correctly, so
+-- QUERY 0 normalises both to an ISO date before filtering. In the live data
+-- every US-format row predates August 2026 anyway (only '12/23/25' and
+-- '12/24/25' occur), so this makes no difference to the current answer --
+-- it's here so the query stays correct if that ever changes.
+-- ============================================================================
+
+
+-- ----------------------------------------------------------------------------
+-- QUERY 0 / THE ANSWER -- one row per qualifying blog, sorted by
+-- earliest_reblog_or_reply_utc then DateAdded descending (normalised --
+-- see the note above). No date cutoff, but scoped to HasBeenOutput = 0
+-- AND IsActive = 1, and now excluding notes on a removed post (see the
+-- header note above). 1,804 rows in the live data as of this revision --
+-- down from 4,396 just before this exclusion was added, because most of
+-- the blogs that dropped out had *every* reblog/reply note pointing at a
+-- now-inactive post, not just some (the number moves between runs
+-- regardless -- crawling and output flip HasBeenOutput/IsActive on live
+-- rows).
+--
+-- earliest_reblog_or_reply_utc is the earliest TimeStamp among this
+-- blog's reblog-or-reply notes (either type counts -- see the column
+-- name); earliest_reblog_or_reply_postid and _root_blogid identify that
+-- specific note's post: PostID + RootBlogId together, not PostID alone --
+-- see TL.db.md ("345 post IDs exist under more than one blog"), same
+-- caution as in find-notes-on-inactive-posts.sql. Resolve RootBlogId to a
+-- name via Blogs (or Blogs, tolerating a miss) if you need it.
+--
+-- Getting "which note" rather than just "when" doesn't fit a plain
+-- MIN()/GROUP BY -- an aggregate can tell you the earliest value but not
+-- which row it came from. EarliestNote instead ranks each blog's
+-- reblog/reply notes with ROW_NUMBER() OVER (PARTITION BY NoteBlogId
+-- ORDER BY TimeStamp), and QUERY 0 takes rn = 1. The ORDER BY carries a
+-- PostID tiebreak because (NoteBlogId, TimeStamp) is not unique in this
+-- data -- ties exist (e.g. NoteBlogId 12 has 7 notes at the same
+-- TimeStamp) -- so without a tiebreak the "earliest" pick would be
+-- arbitrary among ties rather than deterministic.
+--
+-- EarliestNote also excludes notes on an inactive post before ranking
+-- (see the header note above), so "earliest" means earliest surviving
+-- note, not earliest overall -- a blog whose true-earliest note pointed
+-- at a since-removed post now surfaces its next-earliest live one
+-- instead. Applying the exclusion here, not as a filter on QUERY 0's
+-- final rows, matters: filtering after ROW_NUMBER would have picked the
+-- removed-post note as rn = 1 and then dropped the whole row instead of
+-- promoting the next candidate.
+-- ----------------------------------------------------------------------------
+WITH BlogsSplit AS (
+ SELECT
+ b.BlogId,
+ b.BlogName,
+ b.DateAdded,
+ CASE WHEN b.DateAdded LIKE '____-__-__%' THEN 1 ELSE 0 END AS IsIso,
+ -- for the US 'M/d/yy' shape only: everything after the first '/'
+ substr(b.DateAdded, instr(b.DateAdded, '/') + 1) AS RestAfterMonth
+ FROM Blogs b
+ WHERE b.BlogId IS NOT NULL and HasBeenOutput = 0 and IsActive = 1 -- a blog can only match Notes if it has one
+),
+BlogsNorm AS (
+ SELECT
+ BlogId,
+ BlogName,
+ DateAdded,
+ CASE
+ WHEN IsIso = 1 THEN date(DateAdded)
+ ELSE date(
+ '20' || substr(RestAfterMonth, instr(RestAfterMonth, '/') + 1) || '-' ||
+ substr('00' || substr(DateAdded, 1, instr(DateAdded, '/') - 1), -2) || '-' ||
+ substr('00' || substr(RestAfterMonth, 1, instr(RestAfterMonth, '/') - 1), -2)
+ )
+ END AS DateAddedNorm
+ FROM BlogsSplit
+),
+EarliestNote AS (
+ SELECT
+ n.NoteBlogId,
+ n.RootBlogId,
+ n.PostID,
+ n.TimeStamp,
+ ROW_NUMBER() OVER (
+ PARTITION BY n.NoteBlogId
+ ORDER BY n.TimeStamp ASC, n.PostID ASC
+ ) AS rn
+ FROM Notes n
+ JOIN NoteTypes t ON t.TypeId = n.TypeId
+ JOIN Blogs root_bn ON root_bn.BlogId = n.RootBlogId
+ LEFT JOIN Posts p ON p.BlogName = root_bn.BlogName
+ AND p.PostID = n.PostID
+ WHERE t.Type IN ('reblog')--, 'reply')
+ AND n.NoteBlogId IN (SELECT BlogId FROM BlogsNorm) -- scope the window to blogs we care about
+ AND COALESCE(p.IsActive, 1) = 1 -- exclude notes on a post explicitly marked removed
+)
+SELECT
+ bn.BlogId,
+ bn.BlogName || '.tumblr.com',
+ '''' || bn.blogname || ''',',
+ bn.DateAdded,
+ bn.DateAddedNorm,
+ datetime(en.TimeStamp, 'unixepoch') AS earliest_reblog_or_reply_utc,
+ en.PostID AS earliest_reblog_or_reply_postid,
+ en.RootBlogId AS earliest_reblog_or_reply_root_blogid
+ FROM BlogsNorm bn
+ JOIN EarliestNote en ON en.NoteBlogId = bn.BlogId AND en.rn = 1
+ ORDER BY earliest_reblog_or_reply_utc, bn.DateAddedNorm desc
+ limit 50;
+
+
+-- ----------------------------------------------------------------------------
+-- QUERY 1 -- same answer, with a per-blog breakdown of which type(s) fired
+-- and how many. Useful once QUERY 0 has rows; redundant while it's empty.
+-- ----------------------------------------------------------------------------
+-- WITH BlogsSplit AS ( ... ), BlogsNorm AS ( ... ) -- reuse the CTEs above
+--
+-- SELECT
+-- bn.BlogId,
+-- bn.BlogName,
+-- bn.DateAddedNorm,
+-- SUM(CASE WHEN t.Type = 'reblog' THEN 1 ELSE 0 END) AS reblog_count,
+-- SUM(CASE WHEN t.Type = 'reply' THEN 1 ELSE 0 END) AS reply_count
+-- FROM BlogsNorm bn
+-- JOIN Notes n ON n.NoteBlogId = bn.BlogId
+-- JOIN NoteTypes t ON t.TypeId = n.TypeId
+-- WHERE t.Type IN ('reblog', 'reply')
+-- GROUP BY bn.BlogId, bn.BlogName, bn.DateAddedNorm
+-- ORDER BY bn.DateAddedNorm;
+
+
+-- ----------------------------------------------------------------------------
+-- QUERY 2 -- put a date floor back, if wanted later.
+-- Same as QUERY 0, with one extra line in the outer WHERE:
+-- AND bn.DateAddedNorm >= '2026-09-01' -- or whatever cutoff
+-- ----------------------------------------------------------------------------
+
diff --git a/TL.sqbpro b/TL.sqbpro
index 8d0d93d..94606e1 100644
--- a/TL.sqbpro
+++ b/TL.sqbpro
@@ -15,9 +15,9 @@ ORDER BY
FROM
Notes N
inner JOIN
- BlogNames rbn on rbn.BlogId = N.RootBlogId
+ Blogs rbn on rbn.BlogId = N.RootBlogId
inner JOIN
- BlogNames nbn on nbn.BlogId = N.NoteBlogId
+ Blogs nbn on nbn.BlogId = N.NoteBlogId
inner JOIN
NoteTypes t on t.TypeId = N.TypeId
inner JOIN
diff --git a/URLNotesGrabberCORE/DataAccess.cs b/URLNotesGrabberCORE/DataAccess.cs
index 1924966..ad4122a 100644
--- a/URLNotesGrabberCORE/DataAccess.cs
+++ b/URLNotesGrabberCORE/DataAccess.cs
@@ -216,18 +216,19 @@ namespace URLNotesGrabberCORE
#region Notes integer schema
// Notes stopped storing names on 2026-08-07: RootBlogName/NoteBlogName/Type became
- // RootBlogId/NoteBlogId/TypeId, resolved through BlogNames and NoteTypes. There is no
+ // RootBlogId/NoteBlogId/TypeId, resolved through Blogs.BlogId and NoteTypes. There is no
// compatibility view -- a query naming an old column fails outright, so this is a hard
// cut rather than an optional column like IsActive. See TL.db.md.
//
+ // Blogs.BlogId is the only ID authority as of 2026-09-28; the BlogNames table that used
+ // to hold the IDs is gone. Every name in Notes has a Blogs row, created by RegisterBlog.
+ //
// Two shapes recur below and are spelled out inline rather than hidden behind a helper,
// so that every statement reads as the SQL it actually runs:
- // (SELECT BlogId FROM BlogNames WHERE BlogName = @name) -- unique-index probe, 20k rows
+ // (SELECT BlogId FROM Blogs WHERE BlogName = @name) -- primary-key probe
// (SELECT TypeId FROM NoteTypes WHERE Type = 'reply') -- 5 rows, effectively free
- // Joining Notes to Blogs is the one case that must NOT route through BlogNames: Blogs
- // carries its own BlogId, so N.NoteBlogId = B.BlogId is a single integer hop. Joining
- // Notes to Posts is the opposite case -- Posts has only BlogName, so it has to go
- // through BlogNames.
+ // Joining Notes to Blogs is N.NoteBlogId = B.BlogId, a single hop on the unique index.
+ // Joining Notes to Posts goes through Blogs too -- Posts has only BlogName.
///
/// True when the exception is a duplicate-key collision on Notes. The message embeds the
@@ -242,14 +243,30 @@ namespace URLNotesGrabberCORE
}
///
- /// Gives a blog name an ID if it does not have one. No read-back and no round trip -- a
- /// name that is already registered keeps the ID that 1.18M Notes rows point at.
+ /// Ensures a blog has a Blogs row and a BlogId, so a note can point at it. No read-back
+ /// and no round trip -- a blog that already has an ID keeps the one Notes rows point at.
+ /// Unlike AddBlog this does not skip "deact" names: a note by a deactivated blog still
+ /// needs an ID, and Blogs is the only place one can live.
+ /// Assigning the ID is bookkeeping, not a content change, so DateModified is not touched.
+ /// MAX(BlogId) + 1 cannot hand out a used ID because trg_Blogs_BlogId_NoDelete stops any
+ /// row that holds one from being deleted.
///
- private static void RegisterBlogName(SQLiteConnection connection, SQLiteTransaction? transaction, string blogName)
+ private static void RegisterBlog(SQLiteConnection connection, SQLiteTransaction? transaction, string blogName)
{
- using SQLiteCommand command = new SQLiteCommand("INSERT OR IGNORE INTO BlogNames (BlogName) VALUES (@BlogName)", connection, transaction);
- command.Parameters.AddWithValue("@BlogName", blogName);
- command.ExecuteNonQuery();
+ string now = DateTime.Now.ToString("yyyy-MM-dd HH:mm:ss");
+
+ using (SQLiteCommand command = new SQLiteCommand("INSERT OR IGNORE INTO Blogs (BlogName, DateAdded, DateModified, DateCreated) VALUES (@BlogName, @Now, @Now, @Now)", connection, transaction))
+ {
+ command.Parameters.AddWithValue("@BlogName", blogName);
+ command.Parameters.AddWithValue("@Now", now);
+ command.ExecuteNonQuery();
+ }
+
+ using (SQLiteCommand command = new SQLiteCommand("UPDATE Blogs SET BlogId = (SELECT IFNULL(MAX(BlogId), 0) + 1 FROM Blogs) WHERE BlogName = @BlogName AND BlogId IS NULL", connection, transaction))
+ {
+ command.Parameters.AddWithValue("@BlogName", blogName);
+ command.ExecuteNonQuery();
+ }
}
///
@@ -762,7 +779,6 @@ namespace URLNotesGrabberCORE
{
DBPath ??= GetDefaultDbPath();
//try { AddPost(rootBlogName, postID, DBPath); } catch { }
- try { AddBlog(noteBlogName, false, DBPath); } catch { }
using SQLiteConnection connection2 = new SQLiteConnection("Data Source=" + DBPath);
int rowsInserted = 0;
@@ -771,23 +787,24 @@ namespace URLNotesGrabberCORE
{
connection2.Open();
- // Notes stores integer IDs, so both participants and the type have to exist in
- // their lookup table before the note can point at them.
+ // Notes stores integer IDs, so both participants need a Blogs row with a BlogId,
+ // and the type a NoteTypes row, before the note can point at them. RegisterBlog
+ // also does what the AddBlog call here used to: register the note's blog.
//
- // All four statements run in one transaction so a crash cannot leave a name or a
+ // All statements run in one transaction so a crash cannot leave a blog or a
// type registered with no note. The transaction is committed before the console
// output below, which sleeps -- a write lock must not be held across that.
using (SQLiteTransaction transaction = connection2.BeginTransaction())
{
- RegisterBlogName(connection2, transaction, rootBlogName);
- RegisterBlogName(connection2, transaction, noteBlogName);
+ RegisterBlog(connection2, transaction, rootBlogName);
+ RegisterBlog(connection2, transaction, noteBlogName);
RegisterNoteType(connection2, transaction, type ?? string.Empty);
// INSERT OR IGNORE, and no IsActive in the column list: re-crawling a note
// that was removed elsewhere leaves the existing row -- and its flag -- alone.
string sql = "INSERT OR IGNORE INTO Notes (RootBlogId, NoteBlogId, PostID, TimeStamp, TypeId, DatetimeCrawled, DateModified, DateCreated) " +
- "SELECT (SELECT BlogId FROM BlogNames WHERE BlogName = @rootBlogName), " +
- " (SELECT BlogId FROM BlogNames WHERE BlogName = @noteBlogName), " +
+ "SELECT (SELECT BlogId FROM Blogs WHERE BlogName = @rootBlogName), " +
+ " (SELECT BlogId FROM Blogs WHERE BlogName = @noteBlogName), " +
" @PostID, @TimeStamp, " +
" (SELECT TypeId FROM NoteTypes WHERE Type = @Type), " +
" @DatetimeCrawled, @DateModified, @DateCreated";
@@ -1126,11 +1143,11 @@ namespace URLNotesGrabberCORE
{
connection.Open();
- string sql = "SELECT DISTINCT BN.BlogName as blogName, N.PostID" +
+ string sql = "SELECT DISTINCT RB.BlogName as blogName, N.PostID" +
" FROM Notes N" +
- " INNER JOIN BlogNames BN ON BN.BlogId = N.RootBlogId" +
+ " INNER JOIN Blogs RB ON RB.BlogId = N.RootBlogId" +
" WHERE N.TypeId = (SELECT TypeId FROM NoteTypes WHERE Type = 'reply')" + AndIsActive("Notes", "N", DBPath) +
- " ORDER BY BN.BlogName, N.PostID";
+ " ORDER BY RB.BlogName, N.PostID";
using (SQLiteCommand command = new SQLiteCommand(sql, connection))
{
@@ -1171,11 +1188,11 @@ namespace URLNotesGrabberCORE
connection.Open();
// Grouped on the integer rather than the name: the group key is what gets sorted,
- // and BN.BlogName comes along for free off the join.
- string sql = @"SELECT BN.BlogName as blogName, N.PostID,
+ // and RB.BlogName comes along for free off the join.
+ string sql = @"SELECT RB.BlogName as blogName, N.PostID,
MAX(N.TimeStamp) as LatestTimestamp
FROM Notes N
- INNER JOIN BlogNames BN ON BN.BlogId = N.RootBlogId
+ INNER JOIN Blogs RB ON RB.BlogId = N.RootBlogId
WHERE N.TypeId = (SELECT TypeId FROM NoteTypes WHERE Type = 'reply')
AND (N.replyText IS NULL OR N.replyText = '' OR N.replyText = '.')" + AndIsActive("Notes", "N", DBPath) + @"
GROUP BY N.RootBlogId, N.PostID
@@ -1220,13 +1237,13 @@ namespace URLNotesGrabberCORE
{
connection.Open();
- // Posts carries only BlogName, so this is the one join to Notes that has to go
- // through BlogNames -- there is no Posts.BlogId to hop on. The name predicate is
- // pushed into the 20k-row lookup, which then feeds integers to the Notes key.
+ // Posts carries only BlogName, so the join to Notes goes through Blogs -- there is
+ // no Posts.BlogId to hop on. Each post's name is a primary-key probe on Blogs,
+ // which then feeds an integer to the Notes key.
string sql = @"SELECT DISTINCT P.BlogName, P.PostID, MAX(N.TimeStamp) as LatestTimestamp
FROM Posts P
- INNER JOIN BlogNames RBN ON RBN.BlogName = P.BlogName
- INNER JOIN Notes N ON N.RootBlogId = RBN.BlogId AND N.PostID = P.PostID
+ INNER JOIN Blogs RB ON RB.BlogName = P.BlogName
+ INNER JOIN Notes N ON N.RootBlogId = RB.BlogId AND N.PostID = P.PostID
WHERE P.NotFound = 0
AND N.TypeId = (SELECT TypeId FROM NoteTypes WHERE Type = 'reply')
AND (N.replyText IS NULL OR N.replyText = '' OR N.replyText = '.')" + AndIsActive("Posts", "P", DBPath) + AndIsActive("Notes", "N", DBPath) + @"
@@ -1415,8 +1432,7 @@ namespace URLNotesGrabberCORE
try
{
connection.Open();
- // Blogs is reached in one integer hop off Blogs.BlogId, not through BlogNames --
- // that would add a hop and end in the text comparison the migration removed.
+ // Blogs is reached in one integer hop off Blogs.BlogId.
// The negated form is only correct because Notes.TypeId is NOT NULL.
string sql = "";
if (reblogsOnly)
@@ -1703,8 +1719,8 @@ namespace URLNotesGrabberCORE
connection.Open();
string sql = "UPDATE Notes SET TimeStamp = @timestamp, DateModified = @dateModified " +
- "WHERE RootBlogId = (SELECT BlogId FROM BlogNames WHERE BlogName = @rootBlogName) " +
- "AND NoteBlogId = (SELECT BlogId FROM BlogNames WHERE BlogName = @noteBlogName) " +
+ "WHERE RootBlogId = (SELECT BlogId FROM Blogs WHERE BlogName = @rootBlogName) " +
+ "AND NoteBlogId = (SELECT BlogId FROM Blogs WHERE BlogName = @noteBlogName) " +
"AND PostID = @postID AND IFNULL(TimeStamp, 0) <> @timestamp";
using (SQLiteCommand command = new SQLiteCommand(sql, connection))
{
@@ -2018,7 +2034,7 @@ namespace URLNotesGrabberCORE
// Only fan out to rows that match the SELECT criteria in GetRepliesWithFilledText (NULL/empty/legacy-'.'). Never overwrite '?' (confirmed-empty) or already-fetched text.
// The ABS() term cannot use an index on TimeStamp, before or after the integer schema; the NoteBlogId probe is what keeps this off a full scan.
string sql = "UPDATE Notes SET replyText = @replyText, DateModified = @dateModified " +
- "WHERE NoteBlogId = (SELECT BlogId FROM BlogNames WHERE BlogName = @noteBlogName) " +
+ "WHERE NoteBlogId = (SELECT BlogId FROM Blogs WHERE BlogName = @noteBlogName) " +
"AND ABS(TimeStamp - @TimeStamp) <= 5 " +
"AND TypeId = (SELECT TypeId FROM NoteTypes WHERE Type = 'reply') " +
"AND (replyText IS NULL OR replyText = '' OR replyText = '.') " +
@@ -2064,7 +2080,7 @@ namespace URLNotesGrabberCORE
connection.Open();
string sql = "UPDATE Notes SET replyText = @replyText, DateModified = @dateModified " +
- "WHERE RootBlogId = (SELECT BlogId FROM BlogNames WHERE BlogName = @rootBlogName) " +
+ "WHERE RootBlogId = (SELECT BlogId FROM Blogs WHERE BlogName = @rootBlogName) " +
"AND PostID = @PostID " +
"AND TypeId = (SELECT TypeId FROM NoteTypes WHERE Type = 'reply') " +
"AND IFNULL(replyText, '.') <> @replyText";
diff --git a/URLNotesGrabberCORE/TL.db.md b/URLNotesGrabberCORE/TL.db.md
index cc20372..ba695b4 100644
--- a/URLNotesGrabberCORE/TL.db.md
+++ b/URLNotesGrabberCORE/TL.db.md
@@ -24,21 +24,43 @@ Everything below was read out of the live file, not inferred from code. Counts a
> Applied by `../normalize-notes.sql`, which took the file from 207 MB to 148 MB. An
> earlier change the same day (`../shrink-db.sql`) took it from 267 MB to 207 MB.
+> ### ⚠ Breaking change, 2026-09-28: `BlogNames` is gone; `Blogs.BlogId` is the only ID authority
+>
+> The IDs in `Notes` used to live in a `BlogNames` table, with a copy in `Blogs.BlogId`.
+> Nothing kept the copy current, so by 2026-09-28 12,238 blogs first seen after the
+> migration had `Blogs.BlogId = NULL`. Every `Notes`-to-`Blogs` join on `BlogId` silently
+> skipped them and their 23,148 notes, which kept them out of `GetBlogs`.
+>
+> `../retire-blognames.sql` fixed this by giving every note participant a `Blogs` row,
+> backfilling the IDs (none renumbered), making `ix_Blogs_BlogId` unique, and **dropping
+> `BlogNames`**. There is no compatibility view: any query naming it fails with
+> `no such table: BlogNames`. Two triggers now protect the IDs.
+>
+> **Porting an app:** replace `BlogNames` with `Blogs` everywhere. The columns you used,
+> `BlogId` and `BlogName`, exist there with the same meaning. A name lookup
+> (`SELECT BlogId FROM Blogs WHERE BlogName = ?`) is a primary-key probe, and an ID
+> lookup or join (`JOIN Blogs b ON b.BlogId = n.NoteBlogId`) uses the unique
+> `ix_Blogs_BlogId`. Every ID in `Notes` resolves to exactly one `Blogs` row. `Blogs.BlogId`
+> is **no longer** a stale copy, so any code or docs that distrust it can drop that
+> caveat. Never write `BlogId` or `BlogName` on a row that has an ID, and never delete such
+> a row: the triggers reject all three. See [`Blogs`](#blogs).
+
---
## The three content tables
| Table | Rows | What it is |
|---|--:|---|
-| `Blogs` | 188,620 | The crawl registry — one row per known blog, plus crawl-state flags |
+| `Blogs` | 198,560 | The crawl registry, one row per known blog plus crawl-state flags. Also the ID authority for blogs in `Notes` |
| `Posts` | 22,468 | Stored post content. Only 3,867 blogs actually have any |
-| `Notes` | 1,182,333 | The engagement graph: `NoteBlogId` acted on `(RootBlogId, PostID)` |
+| `Notes` | 1,234,830 | The engagement graph: `NoteBlogId` acted on `(RootBlogId, PostID)` |
-…supported by two lookup tables that exist only to keep `Notes` small:
+(`Blogs` and `Notes` counts as of 2026-09-28; the rest as of 2026-08-07.)
+
+…supported by one lookup table that exists only to keep `Notes` small:
| Table | Rows | What it is |
|---|--:|---|
-| `BlogNames` | 20,430 | `BlogId` ⇄ `BlogName`. The ID authority for everything in `Notes` |
| `NoteTypes` | 5 | `TypeId` ⇄ `Type`. `like`, `reblog`, `reply`, `posted`, `post_attribution` |
The engagement graph is the interesting part. 20,311 distinct blogs appear as engagers —
@@ -66,27 +88,47 @@ CREATE TABLE "Blogs" (
PRIMARY KEY("BlogName")
);
-CREATE INDEX ix_Blogs_BlogId ON Blogs (BlogId);
+CREATE UNIQUE INDEX ix_Blogs_BlogId ON Blogs (BlogId);
+
+CREATE TRIGGER trg_Blogs_BlogId_NoDelete -- no DELETE of a row that has a BlogId
+CREATE TRIGGER trg_Blogs_BlogId_Immutable -- no change to its BlogId or BlogName
```
`BlogName` is the primary key, so it is the only indexed way in by name. There is no index
-on any flag or date — filtering or sorting on those scans all 188k rows, which is
+on any flag or date. Filtering or sorting on those scans the whole table, which is
affordable here and is not on `Notes`.
-**`BlogId` is new as of 2026-08-07 and is the join key to `Notes`.** It exists so that
-`Notes` can reach `Blogs` in a single integer hop rather than going through `BlogNames`
-and ending in a text comparison:
+**`BlogId` is the ID that `Notes.RootBlogId` and `Notes.NoteBlogId` store, and `Blogs` is
+the only place it lives** (since 2026-09-28; see the banner at the top). The join to
+`Notes` is one integer hop on the unique index:
```sql
--- what you want
FROM Blogs B JOIN Notes N ON N.NoteBlogId = B.BlogId
-
--- not this
-FROM Blogs B JOIN BlogNames BN ON BN.BlogName = B.BlogName
- JOIN Notes N ON N.NoteBlogId = BN.BlogId
```
-**`BlogId` is NULL on 168,202 of 188,620 rows** — every blog that has never appeared in a
+**Every blog that appears in `Notes` has a `Blogs` row with a `BlogId`.** `AddNote`
+guarantees it through `RegisterBlog`, which runs in the note's own transaction:
+
+```sql
+INSERT OR IGNORE INTO Blogs (BlogName, DateAdded, DateModified, DateCreated)
+VALUES (@name, @now, @now, @now);
+UPDATE Blogs SET BlogId = (SELECT IFNULL(MAX(BlogId), 0) + 1 FROM Blogs)
+ WHERE BlogName = @name AND BlogId IS NULL;
+```
+
+- Unlike `AddBlog`, this does **not** skip names containing `deact`. A note by a
+ deactivated blog still needs an ID, so such blogs now get registry rows too, with the
+ usual defaults (`HasBeenOutput = 0`, `IsActive` left at its default).
+- Assigning a `BlogId` is bookkeeping, so it **does not move `DateModified`**.
+- `MAX(BlogId) + 1` is safe only because an ID can never be freed. The two triggers see
+ to that: deleting a row that has a `BlogId`, or changing its `BlogId` or `BlogName`,
+ aborts. Remove a blog with `IsActive = 0` instead. A blog renamed upstream gets a new
+ row. Rows with no `BlogId` can still be deleted or renamed freely.
+- `INSERT OR REPLACE` on `Blogs` gets around the delete trigger (SQLite does not fire
+ delete triggers for REPLACE unless `recursive_triggers` is on), and it would wipe the
+ `BlogId`. It was already forbidden because it resets `IsActive`. Do not use it.
+
+**`BlogId` is NULL on 165,887 of 198,560 rows**, every blog that has never appeared in a
note. That is the large majority, and it is not an error: the registry is far bigger than
the engagement graph. An inner join on `BlogId` therefore silently drops those blogs,
which is usually what you want for engagement queries and is wrong for registry listings.
@@ -187,8 +229,7 @@ CREATE INDEX ix_Notes_NoteBlogId ON Notes (NoteBlogId);
**Integer IDs since 2026-08-07 — this is the breaking change.** `RootBlogName`,
`NoteBlogName` and `Type` are gone, replaced by `RootBlogId`, `NoteBlogId` and `TypeId`.
-Resolve them through [`BlogNames`](#blognames) and [`NoteTypes`](#notetypes), or join
-straight to `Blogs` on `BlogId`. The old names were text repeated on 1.18 million rows,
+Resolve blog IDs through `Blogs.BlogId` and types through [`NoteTypes`](#notetypes). The old names were text repeated on 1.18 million rows,
in the table *and* in every index over it; the swap took the file from 207 MB to 148 MB.
The **primary key column order is deliberately unchanged**, so the leading-prefix access
@@ -251,43 +292,21 @@ At 1.18M rows this is the table that dictates how the whole database has to be q
- `replyText` is `'.'` on 1,167,464 rows — only `reply` notes carry real text. Those
dots are inherited from the old column default; new rows get `NULL` instead.
-**Resolve IDs by filtering the lookup, not by scanning `Notes`.** The lookup tables are
-tiny and uniquely indexed, so pushing a name predicate into them costs nothing and lets
-the `Notes` index do the work:
+**Resolve IDs by filtering `Blogs`, not by scanning `Notes`.** A name predicate on `Blogs`
+is a primary-key probe, so pushing it there costs nothing and lets the `Notes` index do
+the work:
```sql
--- good: BlogNames resolves the name, then the index is searched
+-- good: Blogs resolves the name, then the index is searched
SELECT * FROM Notes
- WHERE NoteBlogId = (SELECT BlogId FROM BlogNames WHERE BlogName = ?);
+ WHERE NoteBlogId = (SELECT BlogId FROM Blogs WHERE BlogName = ?);
-- also good, same plan
SELECT n.* FROM Notes n
- JOIN BlogNames b ON b.BlogId = n.NoteBlogId
+ JOIN Blogs b ON b.BlogId = n.NoteBlogId
WHERE b.BlogName = ?;
```
-### `BlogNames`
-
-```sql
-CREATE TABLE BlogNames (
- BlogId INTEGER PRIMARY KEY,
- BlogName TEXT NOT NULL UNIQUE
-);
-```
-
-20,430 rows — every name appearing in `Notes` as either participant, and nothing else.
-This is the **ID authority**: `Notes.RootBlogId` and `Notes.NoteBlogId` both point here,
-and `Blogs.BlogId` is a copy of the value for the blogs that have one.
-
-**12 of these names have no `Blogs` row.** The registry has never been a superset of the
-engagement graph and still is not, so resolving an ID through `Blogs` rather than
-`BlogNames` will occasionally find nothing. Use `BlogNames` when you need the name itself
-and `Blogs` when you need registry columns.
-
-IDs are assigned by SQLite and are **stable**: they are stored in over a million `Notes`
-rows. Never renumber them. A blog that is renamed upstream should get a new row, not an
-edit to an existing one, unless every `Notes` reference is migrated with it.
-
### `NoteTypes`
```sql
@@ -317,6 +336,9 @@ code to this table's contents, so prefer the join in anything long-lived.
## Porting to the integer schema
+> Written for the 2026-08-07 change, and updated for 2026-09-28: wherever this section
+> once said `BlogNames`, it now says `Blogs`. `BlogNames` no longer exists.
+
Everything here was checked against the live 148 MB file. There were 14 affected call
sites in `DataAccess.cs` and 16 in `RolodexRepository.cs`. TumblThree needs no changes —
its single statement touches `Blogs.IsActive` and `BlogName` only.
@@ -333,8 +355,8 @@ the result.
| Was | Is now | Resolve via |
|---|---|---|
-| `Notes.RootBlogName` | `Notes.RootBlogId` | `BlogNames.BlogId` → `.BlogName` |
-| `Notes.NoteBlogName` | `Notes.NoteBlogId` | `BlogNames.BlogId` → `.BlogName` |
+| `Notes.RootBlogName` | `Notes.RootBlogId` | `Blogs.BlogId` → `.BlogName` |
+| `Notes.NoteBlogName` | `Notes.NoteBlogId` | `Blogs.BlogId` → `.BlogName` |
| `Notes.Type` | `Notes.TypeId` | `NoteTypes.TypeId` → `.Type` |
| `ix_NoteBlogName01` | `ix_Notes_NoteBlogId` | — |
@@ -347,14 +369,14 @@ the result.
-- was
WHERE NoteBlogName = @Name
--- now, either form; both search ix_Notes_NoteBlogId after a unique-index lookup
-WHERE NoteBlogId = (SELECT BlogId FROM BlogNames WHERE BlogName = @Name)
+-- now, either form; both search ix_Notes_NoteBlogId after a primary-key lookup
+WHERE NoteBlogId = (SELECT BlogId FROM Blogs WHERE BlogName = @Name)
-- or
-JOIN BlogNames b ON b.BlogId = n.NoteBlogId WHERE b.BlogName = @Name
+JOIN Blogs b ON b.BlogId = n.NoteBlogId WHERE b.BlogName = @Name
```
-Measured 73 ms against 63 ms for the old text form on the busiest blog — the extra hop is
-a unique-index probe on a 20k-row table and does not show.
+Measured at 73 ms against 63 ms for the old text form on the busiest blog, when the lookup
+was still `BlogNames`. The extra hop is one index probe and does not show.
### Joining `Notes` to `Blogs`
@@ -364,12 +386,17 @@ This is the join to get right; it is the most common shape in both applications.
-- was
FROM Blogs B INNER JOIN Notes N ON N.NoteBlogName = B.BlogName
--- now: one integer hop, using the new Blogs.BlogId
+-- now: one integer hop, using Blogs.BlogId
FROM Blogs B INNER JOIN Notes N ON N.NoteBlogId = B.BlogId
```
-Do **not** route this through `BlogNames` — that adds a hop and ends in the text
-comparison the change was meant to remove.
+Joining `Notes` to `Posts` also goes through `Blogs`, since `Posts` has only a name:
+
+```sql
+FROM Posts P
+JOIN Blogs RB ON RB.BlogName = P.BlogName
+JOIN Notes N ON N.RootBlogId = RB.BlogId AND N.PostID = P.PostID
+```
### Selecting a name back out
@@ -378,12 +405,12 @@ comparison the change was meant to remove.
SELECT NoteBlogName AS blogName, COUNT(*) FROM Notes ... GROUP BY NoteBlogName
-- now
-SELECT bn.BlogName AS blogName, COUNT(*)
- FROM Notes n JOIN BlogNames bn ON bn.BlogId = n.NoteBlogId
- ... GROUP BY bn.BlogName
+SELECT b.BlogName AS blogName, COUNT(*)
+ FROM Notes n JOIN Blogs b ON b.BlogId = n.NoteBlogId
+ ... GROUP BY b.BlogName
```
-Group by `n.NoteBlogId` instead of `bn.BlogName` when you only need the name for display —
+Group by `n.NoteBlogId` instead of `b.BlogName` when you only need the name for display —
grouping on the integer is cheaper and the name comes along for free.
### Filtering by type
@@ -407,27 +434,25 @@ because `TypeId` is `NOT NULL`.
### Inserting a note
-The crawler must ensure both names have IDs first. `INSERT OR IGNORE` on `BlogNames` is
-the whole of it — no read-back, no round trip, safe to run every time:
+The crawler must ensure both blogs have IDs first: run the `RegisterBlog` pair shown under
+[`Blogs`](#blogs) for each name. No read-back, no round trip, and safe to run every time.
+Then:
```sql
-INSERT OR IGNORE INTO BlogNames (BlogName) VALUES (@rootBlogName);
-INSERT OR IGNORE INTO BlogNames (BlogName) VALUES (@noteBlogName);
-
INSERT OR IGNORE INTO Notes
(RootBlogId, PostID, NoteBlogId, TimeStamp, TypeId,
DatetimeCrawled, DateModified, DateCreated)
-SELECT (SELECT BlogId FROM BlogNames WHERE BlogName = @rootBlogName),
+SELECT (SELECT BlogId FROM Blogs WHERE BlogName = @rootBlogName),
@PostID,
- (SELECT BlogId FROM BlogNames WHERE BlogName = @noteBlogName),
+ (SELECT BlogId FROM Blogs WHERE BlogName = @noteBlogName),
@TimeStamp,
(SELECT TypeId FROM NoteTypes WHERE Type = @Type),
@DatetimeCrawled, @DateModified, @DateCreated;
```
Verified: a genuinely new note inserts, and re-running the identical statement inserts 0.
-Run all three statements in one transaction so a crash cannot leave a name registered
-with no note.
+Run the registrations and the insert in one transaction so a crash cannot leave a blog
+registered with no note.
**The duplicate-key error message has changed.** `DataAccess.cs` compares against the
literal string
@@ -452,7 +477,7 @@ on `TimeStamp` either before or after:
```sql
-- now
UPDATE Notes SET replyText = @replyText, DateModified = @dateModified
- WHERE NoteBlogId = (SELECT BlogId FROM BlogNames WHERE BlogName = @noteBlogName)
+ WHERE NoteBlogId = (SELECT BlogId FROM Blogs WHERE BlogName = @noteBlogName)
AND ABS(TimeStamp - @TimeStamp) <= 5
AND TypeId = (SELECT TypeId FROM NoteTypes WHERE Type = 'reply')
AND (replyText IS NULL OR replyText = '' OR replyText = '.')
@@ -462,18 +487,15 @@ UPDATE Notes SET replyText = @replyText, DateModified = @dateModified
Rolodex's soft-delete updates need no change beyond the `WHERE` clause — they set
`IsActive`, which is untouched.
-### Three traps
+### Two traps
-**`Blogs.BlogId` is NULL on 168,202 of 188,620 rows.** Any inner join on it silently drops
+**`Blogs.BlogId` is NULL on 165,887 of 198,560 rows.** Any inner join on it silently drops
every blog that has never appeared in a note. Correct for engagement queries; wrong for
registry listings, which need a `LEFT JOIN` or no join at all.
-**12 names in `BlogNames` have no `Blogs` row.** Resolving an ID to a name through `Blogs`
-will occasionally find nothing. Use `BlogNames` for names and `Blogs` for registry columns.
-
-**IDs are stable and must stay so.** `BlogNames.BlogId` and `NoteTypes.TypeId` are stored
-in over a million `Notes` rows. Never renumber. A blog renamed upstream gets a new row,
-not an edited one, unless every `Notes` reference migrates with it.
+**IDs are stable and must stay so.** `Blogs.BlogId` and `NoteTypes.TypeId` are stored in
+over a million `Notes` rows. Never renumber. A blog renamed upstream gets a new row, not
+an edited one. The `Blogs` triggers reject both.
---
@@ -481,16 +503,12 @@ not an edited one, unless every `Notes` reference migrates with it.
There are no foreign keys, and the tables do not perfectly agree:
-- 4 `Posts` rows name a blog with no `Blogs` row.
-- 12 of the 20,430 names in `BlogNames` have no `Blogs` row.
-
-So a name appearing in `Notes` or `Posts` is not a guarantee that the registry knows about
-it. Joins from those tables back to `Blogs` should tolerate a miss.
-
-The integer schema does not fix this and was not meant to. `BlogNames` is deliberately
-built from `Notes` rather than from `Blogs`, precisely so that the 12 unregistered
-engagers keep their IDs and their rows. Had it been built from the registry, those notes
-would have been dropped by the migration's inner joins.
+- 4 `Posts` rows name a blog with no `Blogs` row, so joins from `Posts` back to `Blogs`
+ should tolerate a miss.
+- `Notes` is covered: every `RootBlogId` and `NoteBlogId` resolves to a `Blogs` row.
+ `retire-blognames.sql` checked this before committing, and `RegisterBlog` keeps it true.
+ Before 2026-09-28, 12 to 17 note participants had no registry row. They now have stub
+ rows.
---
@@ -542,11 +560,10 @@ Crawler bookkeeping. Rolodex ignores all of these.
`DataAccess.cs` joins on it to decide what to collect:
```sql
--- shape only; the ported GetBlogs joins Blogs directly on BlogId and needs no BlogNames hop
-SELECT bn.BlogName, count(*)
+-- shape only
+SELECT b.BlogName, count(*)
FROM Notes n
- JOIN Blogs b ON b.BlogId = n.NoteBlogId
- JOIN BlogNames bn ON bn.BlogId = n.NoteBlogId
+ JOIN Blogs b ON b.BlogId = n.NoteBlogId
WHERE b.IsActive = @isActive AND ...
```
@@ -644,7 +661,7 @@ handled:
SELECT 'Blogs', COUNT(*) FROM Blogs
UNION ALL SELECT 'Posts', COUNT(*) FROM Posts
UNION ALL SELECT 'Notes', COUNT(*) FROM Notes
-UNION ALL SELECT 'BlogNames', COUNT(*) FROM BlogNames;
+UNION ALL SELECT 'Blogs with a BlogId', COUNT(*) FROM Blogs WHERE BlogId IS NOT NULL;
-- note type mix (joins NoteTypes; Notes.Type no longer exists)
SELECT t.Type, COUNT(*)
@@ -668,8 +685,9 @@ SELECT COUNT(*) FROM (
SELECT COUNT(*) FROM Posts p
WHERE NOT EXISTS (SELECT 1 FROM Blogs b WHERE b.BlogName = p.BlogName);
-SELECT COUNT(*) FROM BlogNames bn
-WHERE NOT EXISTS (SELECT 1 FROM Blogs b WHERE b.BlogName = bn.BlogName);
+-- note participants with no Blogs.BlogId (expect 0; anything else is the pre-2026-09-28 drift)
+SELECT COUNT(*) FROM (SELECT DISTINCT NoteBlogId AS Id FROM Notes) n
+WHERE NOT EXISTS (SELECT 1 FROM Blogs b WHERE b.BlogId = n.Id);
-- space by object, to see where the file actually goes
SELECT name, SUM(pgsize)/1024/1024 AS mb
diff --git a/normalize-notes.sql b/normalize-notes.sql
index 5602564..c5db9d7 100644
--- a/normalize-notes.sql
+++ b/normalize-notes.sql
@@ -2,6 +2,10 @@
-- Replaces the repeated blog-name and type TEXT in Notes with integer IDs.
-- Reduces TL.db from ~207 MB to ~148 MB (-29%).
--
+-- SUPERSEDED IN PART, 2026-09-28: the BlogNames table this creates is no longer the
+-- ID authority. Run retire-blognames.sql straight after this one; it moves the IDs
+-- into Blogs.BlogId and drops BlogNames. The current app code assumes both have run.
+--
-- THIS IS A BREAKING SCHEMA CHANGE. There is no compatibility layer. Every
-- query in URLNotesGrabberCORE and Rolodex that names Notes.RootBlogName,
-- Notes.NoteBlogName or Notes.Type stops working the moment this runs, and
diff --git a/retire-blognames.sql b/retire-blognames.sql
new file mode 100644
index 0000000..8d0781f
--- /dev/null
+++ b/retire-blognames.sql
@@ -0,0 +1,143 @@
+-- retire-blognames.sql
+-- Makes Blogs.BlogId the only ID authority for Notes and retires the BlogNames table.
+--
+-- WHY: normalize-notes.sql (2026-08-07) put the IDs in BlogNames and copied them into
+-- Blogs.BlogId once. Nothing kept the copy current: by 2026-09-28, 12,238 blogs first seen
+-- in a note after the migration had a BlogNames ID but Blogs.BlogId = NULL, so every query
+-- joining Notes to Blogs on BlogId (GetBlogs and friends) silently skipped them -- 23,148
+-- notes. Two copies of one ID drift; this leaves one.
+--
+-- What it does:
+-- 1. Gives every BlogNames name a Blogs row (17 had none), carrying its ID over.
+-- 2. Copies the ID onto every Blogs row that is missing it. IDs are never renumbered --
+-- they are stored in 1.18M Notes rows.
+-- 3. Proves every Notes ID resolves through Blogs before anything is dropped.
+-- 4. Makes ix_Blogs_BlogId UNIQUE.
+-- 5. Drops BlogNames. No compatibility view: any other app that still names it gets
+-- "no such table: BlogNames" and must port to Blogs.BlogId (see TL.db.md).
+-- 6. Adds triggers that stop a Blogs row holding a BlogId from being deleted, renamed or
+-- renumbered -- the guarantees BlogNames gave by never being touched.
+--
+-- DateModified is NOT moved: assigning an ID is bookkeeping, not a content change. The 17
+-- new stub rows get DateAdded/DateModified/DateCreated = now, as AddBlog would give them.
+--
+-- Runs after normalize-notes.sql. A backup from before 2026-08-07 needs both, in order.
+--
+-- HOW TO RUN:
+-- 1. Stop every app that uses TL.db. Pause NextCloud sync.
+-- 2. Back up TL.db: sqlite3 TL.db ".backup 'TL pre-retire-blognames.db'"
+-- 3. sqlite3 -bail TL.db < retire-blognames.sql
+-- -bail matters: a failed check aborts before COMMIT and nothing is changed.
+-- In DB Browser, Execute SQL stops at the first error; then Revert Changes.
+-- 4. Run the build of URLNotesGrabberCORE that no longer uses BlogNames. An older
+-- build fails every AddNote with "no such table: BlogNames".
+
+PRAGMA foreign_keys = off;
+
+BEGIN;
+
+-- Every check inserts one count here; the CHECK aborts the script on anything but 0.
+CREATE TEMP TABLE MustBeZero (Check_ TEXT, n INTEGER CHECK (n = 0));
+
+--------------------------------------------------------------------------
+-- STEP 0: the two copies must not disagree anywhere they are both set
+--------------------------------------------------------------------------
+INSERT INTO MustBeZero
+SELECT 'Blogs.BlogId differs from BlogNames', COUNT(*)
+ FROM Blogs b JOIN BlogNames bn ON bn.BlogName = b.BlogName
+ WHERE b.BlogId <> bn.BlogId;
+
+INSERT INTO MustBeZero
+SELECT 'Blogs.BlogId unknown to BlogNames', COUNT(*)
+ FROM Blogs b
+ WHERE b.BlogId IS NOT NULL
+ AND NOT EXISTS (SELECT 1 FROM BlogNames bn WHERE bn.BlogId = b.BlogId AND bn.BlogName = b.BlogName);
+
+--------------------------------------------------------------------------
+-- STEP 1: a Blogs row for every name Notes points at
+--------------------------------------------------------------------------
+-- No IsActive in the column list: it is not ours to write (defaults to live).
+INSERT INTO Blogs (BlogName, DateAdded, DateModified, DateCreated, BlogId)
+SELECT bn.BlogName,
+ strftime('%Y-%m-%d %H:%M:%S', 'now', 'localtime'),
+ strftime('%Y-%m-%d %H:%M:%S', 'now', 'localtime'),
+ strftime('%Y-%m-%d %H:%M:%S', 'now', 'localtime'),
+ bn.BlogId
+ FROM BlogNames bn
+ WHERE NOT EXISTS (SELECT 1 FROM Blogs b WHERE b.BlogName = bn.BlogName);
+
+--------------------------------------------------------------------------
+-- STEP 2: backfill the IDs Blogs never received
+--------------------------------------------------------------------------
+UPDATE Blogs
+ SET BlogId = (SELECT bn.BlogId FROM BlogNames bn WHERE bn.BlogName = Blogs.BlogName)
+ WHERE BlogId IS NULL
+ AND BlogName IN (SELECT BlogName FROM BlogNames);
+
+--------------------------------------------------------------------------
+-- STEP 3: prove Blogs now holds exactly what BlogNames held
+--------------------------------------------------------------------------
+INSERT INTO MustBeZero
+SELECT 'BlogNames pair missing from Blogs', COUNT(*)
+ FROM BlogNames bn
+ WHERE NOT EXISTS (SELECT 1 FROM Blogs b WHERE b.BlogId = bn.BlogId AND b.BlogName = bn.BlogName);
+
+INSERT INTO MustBeZero
+SELECT 'Blogs IDs vs BlogNames rows', (SELECT COUNT(*) FROM Blogs WHERE BlogId IS NOT NULL) - (SELECT COUNT(*) FROM BlogNames);
+
+INSERT INTO MustBeZero
+SELECT 'Notes.RootBlogId unresolved', COUNT(*)
+ FROM (SELECT DISTINCT RootBlogId AS Id FROM Notes) n
+ WHERE NOT EXISTS (SELECT 1 FROM Blogs b WHERE b.BlogId = n.Id);
+
+INSERT INTO MustBeZero
+SELECT 'Notes.NoteBlogId unresolved', COUNT(*)
+ FROM (SELECT DISTINCT NoteBlogId AS Id FROM Notes) n
+ WHERE NOT EXISTS (SELECT 1 FROM Blogs b WHERE b.BlogId = n.Id);
+
+--------------------------------------------------------------------------
+-- STEP 4: one row per ID
+--------------------------------------------------------------------------
+-- UNIQUE still allows the NULLs on the ~168k blogs that have never appeared in a note.
+DROP INDEX ix_Blogs_BlogId;
+CREATE UNIQUE INDEX ix_Blogs_BlogId ON Blogs (BlogId);
+
+--------------------------------------------------------------------------
+-- STEP 5: BlogNames goes
+--------------------------------------------------------------------------
+DROP TABLE BlogNames;
+
+--------------------------------------------------------------------------
+-- STEP 6: what BlogNames guaranteed by never being written
+--------------------------------------------------------------------------
+-- A deleted row would orphan its notes, and MAX(BlogId) + 1 in RegisterBlog could then
+-- hand the same ID to a different blog. Remove a blog with IsActive = 0 instead.
+CREATE TRIGGER trg_Blogs_BlogId_NoDelete
+BEFORE DELETE ON Blogs
+WHEN OLD.BlogId IS NOT NULL
+BEGIN
+ SELECT RAISE(ABORT, 'Blogs row has a BlogId that Notes points at; set IsActive = 0 instead of deleting');
+END;
+
+-- A blog renamed upstream is a new blog to Tumblr's API and gets a new row. Editing the name
+-- in place would re-attribute every note to it; changing the ID would orphan them.
+CREATE TRIGGER trg_Blogs_BlogId_Immutable
+BEFORE UPDATE OF BlogId, BlogName ON Blogs
+WHEN OLD.BlogId IS NOT NULL
+ AND (NEW.BlogId IS NOT OLD.BlogId OR NEW.BlogName IS NOT OLD.BlogName)
+BEGIN
+ SELECT RAISE(ABORT, 'BlogId and BlogName are fixed once a blog has a BlogId; Notes rows point at it');
+END;
+
+DROP TABLE temp.MustBeZero;
+
+COMMIT;
+
+--------------------------------------------------------------------------
+-- VERIFY
+--------------------------------------------------------------------------
+-- SELECT COUNT(*) FROM sqlite_master WHERE name = 'BlogNames'; -- expect: 0
+-- SELECT COUNT(*) FROM Blogs WHERE BlogId IS NOT NULL; -- expect: the old BlogNames row count
+-- SELECT sql FROM sqlite_master WHERE name = 'ix_Blogs_BlogId'; -- expect: CREATE UNIQUE INDEX
+-- SELECT name FROM sqlite_master WHERE type = 'trigger'; -- expect: both triggers
+-- PRAGMA integrity_check; -- expect: ok
diff --git a/verify-db-schema.sql b/verify-db-schema.sql
index 43a4dea..62b533a 100644
--- a/verify-db-schema.sql
+++ b/verify-db-schema.sql
@@ -82,7 +82,8 @@ WITH expected(tbl, col, alter_stmt) AS (
-- Blogs.BlogId (2026-08-07) is the single-hop join key into Notes. Deliberately NOT
-- auto-fixable: an added-but-empty BlogId makes every engagement join return zero
-- rows silently, which is worse than the hard error a missing column gives.
- ('Blogs','BlogId', 'MANUAL REVIEW - see query 1d: run normalize-notes.sql'),
+ -- Since 2026-09-28 it is the only blog-ID authority (query 1e).
+ ('Blogs','BlogId', 'MANUAL REVIEW - see queries 1d/1e: run normalize-notes.sql, then retire-blognames.sql'),
-- Notes (base columns: manual review if missing)
-- Integer IDs since 2026-08-07. RootBlogName/NoteBlogName/Type are GONE, not renamed
@@ -101,11 +102,11 @@ WITH expected(tbl, col, alter_stmt) AS (
-- placeholder. EnsureReplyTextColumnExists in DataAccess.cs adds it the same way.
('Notes','replyText', 'ALTER TABLE Notes ADD COLUMN replyText TEXT;'),
- -- BlogNames / NoteTypes (the lookup tables Notes resolves its IDs through, 2026-08-07).
- -- Not auto-fixable: an empty BlogNames does not mean "add the table", it means the
+ -- NoteTypes (the lookup table Notes resolves TypeId through, 2026-08-07).
+ -- Not auto-fixable: an empty NoteTypes does not mean "add the table", it means the
-- Notes rows have nothing to resolve against. Rebuild with normalize-notes.sql.
- ('BlogNames','BlogId', 'MANUAL REVIEW - see query 1d: run normalize-notes.sql'),
- ('BlogNames','BlogName', 'MANUAL REVIEW - see query 1d: run normalize-notes.sql'),
+ -- BlogNames is not listed: it was dropped on 2026-09-28. Query 1e reports a file
+ -- that still has it.
('NoteTypes','TypeId', 'MANUAL REVIEW - see query 1d: run normalize-notes.sql'),
('NoteTypes','Type', 'MANUAL REVIEW - see query 1d: run normalize-notes.sql'),
@@ -123,7 +124,6 @@ actual(tbl, col) AS (
SELECT 'Posts', name FROM pragma_table_info('Posts')
UNION ALL SELECT 'Blogs', name FROM pragma_table_info('Blogs')
UNION ALL SELECT 'Notes', name FROM pragma_table_info('Notes')
- UNION ALL SELECT 'BlogNames', name FROM pragma_table_info('BlogNames')
UNION ALL SELECT 'NoteTypes', name FROM pragma_table_info('NoteTypes')
UNION ALL SELECT 'DailyAPICount', name FROM pragma_table_info('DailyAPICount')
UNION ALL SELECT 'ApiKeyPoolState', name FROM pragma_table_info('ApiKeyPoolState')
@@ -146,7 +146,7 @@ ORDER BY (e.alter_stmt LIKE 'ALTER%') DESC, e.tbl, e.col;
-- 1b. MISSING TABLES: expected tables that don't exist at all in this DB.
-- Zero rows = good.
WITH expected_tables(tbl) AS (
- VALUES ('Posts'),('Blogs'),('Notes'),('BlogNames'),('NoteTypes'),('DailyAPICount'),
+ VALUES ('Posts'),('Blogs'),('Notes'),('NoteTypes'),('DailyAPICount'),
('ApiKeyPoolState'),('ApiKeyPoolMeta')
)
SELECT et.tbl AS missing_table
@@ -180,7 +180,6 @@ WITH expected(tbl, col) AS (
('Notes','RootBlogId'),('Notes','PostID'),('Notes','NoteBlogId'),('Notes','TimeStamp'),
('Notes','TypeId'),('Notes','DatetimeCrawled'),('Notes','DateModified'),('Notes','DateCreated'),
('Notes','replyText'),('Notes','IsActive'),
- ('BlogNames','BlogId'),('BlogNames','BlogName'),
('NoteTypes','TypeId'),('NoteTypes','Type'),
('DailyAPICount','Date'),('DailyAPICount','APICount'),
('ApiKeyPoolState','KeyName'),('ApiKeyPoolState','RetryUntil'),
@@ -190,7 +189,6 @@ actual(tbl, col) AS (
SELECT 'Posts', name FROM pragma_table_info('Posts')
UNION ALL SELECT 'Blogs', name FROM pragma_table_info('Blogs')
UNION ALL SELECT 'Notes', name FROM pragma_table_info('Notes')
- UNION ALL SELECT 'BlogNames', name FROM pragma_table_info('BlogNames')
UNION ALL SELECT 'NoteTypes', name FROM pragma_table_info('NoteTypes')
UNION ALL SELECT 'DailyAPICount', name FROM pragma_table_info('DailyAPICount')
UNION ALL SELECT 'ApiKeyPoolState', name FROM pragma_table_info('ApiKeyPoolState')
@@ -209,7 +207,7 @@ ORDER BY a.tbl, a.col;
--
-- This is the one failure SECTION 2 cannot fix. Notes.RootBlogName /
-- NoteBlogName / Type were replaced by RootBlogId / NoteBlogId / TypeId
--- resolving through BlogNames and NoteTypes -- a data migration, not an
+-- resolving through (then) BlogNames and NoteTypes -- a data migration, not an
-- ADD COLUMN. There is no compatibility view, so the current code fails
-- outright ("no such column: RootBlogId") against such a file.
--
@@ -223,6 +221,28 @@ WHERE lower(name) IN ('rootblogname','noteblogname','type')
HAVING COUNT(*) > 0;
+-- 1e. BLOGNAMES NOT RETIRED: a backup from between 2026-08-07 and 2026-09-28, when
+-- BlogNames still held the IDs and Blogs.BlogId was an
+-- unmaintained copy. Zero rows = good.
+--
+-- The current code resolves every Notes ID through Blogs.BlogId and never writes
+-- BlogNames, so against such a file new blogs get IDs that can collide with
+-- BlogNames' and every blog missing from Blogs.BlogId stays invisible to GetBlogs.
+--
+-- Fix: back up, then run retire-blognames.sql (after normalize-notes.sql if 1d
+-- also reported). It checks itself and changes nothing if a check fails.
+SELECT 'BlogNames still exists (' || type || ') -- run retire-blognames.sql' AS blognames_not_retired
+FROM sqlite_master
+WHERE lower(name) = 'blognames'
+UNION ALL
+SELECT 'Blogs.BlogId is not UNIQUE -- run retire-blognames.sql'
+WHERE NOT EXISTS (SELECT 1 FROM pragma_index_list('Blogs') WHERE name = 'ix_Blogs_BlogId' AND "unique" = 1)
+UNION ALL
+SELECT 'BlogId guard trigger missing: ' || t.name || ' -- run retire-blognames.sql'
+FROM (SELECT 'trg_Blogs_BlogId_NoDelete' AS name UNION ALL SELECT 'trg_Blogs_BlogId_Immutable') t
+WHERE NOT EXISTS (SELECT 1 FROM sqlite_master m WHERE m.type = 'trigger' AND m.name = t.name);
+
+
-- ============================================================================
-- SECTION 2 -- FIX (opt-in, additive only)
--
@@ -233,8 +253,8 @@ HAVING COUNT(*) > 0;
-- subset. These are the 8 additive migration columns and nothing else; the
-- likes high-water-mark reset is intentionally NOT included.
--
--- Nothing here addresses query 1d. The Notes integer schema is a data migration
--- (normalize-notes.sql) and cannot be reached by adding columns.
+-- Nothing here addresses queries 1d or 1e. Those are data migrations
+-- (normalize-notes.sql, retire-blognames.sql) and cannot be reached by adding columns.
-- ============================================================================
-- ALTER TABLE Posts ADD COLUMN PostType TEXT;