diff --git a/AGENTS.md b/AGENTS.md index 066e933..04445fc 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -49,28 +49,40 @@ say nothing about the item being fetched, so they must not be recorded as per-it ### `Notes` Stores Integer IDs, Not Names As of 2026-08-07 `Notes.RootBlogName`, `NoteBlogName` and `Type` are gone, replaced by -`RootBlogId`, `NoteBlogId` and `TypeId` resolving through the `BlogNames` and `NoteTypes` -lookup tables. There is no compatibility view — naming an old column is a hard SQLite -error, so unlike `IsActive` this is a hard cut with no runtime probe. Full detail in -`URLNotesGrabberCORE/TL.db.md`. +`RootBlogId`, `NoteBlogId` and `TypeId`. Blog IDs resolve through `Blogs.BlogId`, and +types through the `NoteTypes` lookup table. There is no compatibility view: naming an old +column is a hard SQLite error, so unlike `IsActive` this is a hard cut with no runtime +probe. Full detail in `URLNotesGrabberCORE/TL.db.md`. -- **Joining `Notes` to `Blogs` goes through `Blogs.BlogId`**, not `BlogNames`: - `FROM Blogs B INNER JOIN Notes N ON N.NoteBlogId = B.BlogId`. Routing it through - `BlogNames` adds a hop and ends in the text comparison the migration removed -- **Joining `Notes` to `Posts` is the opposite** — `Posts` has only `BlogName`, so it must - go through `BlogNames` (`GetRepliesWithFilledText`). This is the only such join -- **Resolve a name by filtering the lookup, never by scanning `Notes`**: - `WHERE NoteBlogId = (SELECT BlogId FROM BlogNames WHERE BlogName = @name)`. The subquery - is a unique-index probe on 20k rows and does not show against the 1.18M-row table -- **`AddNote` registers both blog names *and* the note type** with `INSERT OR IGNORE` - before inserting, all in one transaction. `NoteTypes` is a table rather than a `CHECK` - constraint precisely so an unseen type is an `INSERT`; without that registration it - would resolve to `NULL` and fail the `NOT NULL` on `TypeId`, losing the note -- **`Blogs.BlogId` is NULL on 168,202 of 188,620 rows** — every blog that has never - appeared in a note. An inner join on it silently drops them. Correct for engagement - queries, wrong for anything listing the registry -- **IDs are stable and must never be renumbered.** They are stored in 1.18M `Notes` rows. - A blog renamed upstream gets a new `BlogNames` row, not an edited one +- **`Blogs.BlogId` is the only blog-ID authority (since 2026-09-28).** IDs used to live in a + `BlogNames` table with an unmaintained copy in `Blogs.BlogId`. The copy drifted and hid + 12k blogs from `GetBlogs`, so `retire-blognames.sql` moved the authority into `Blogs` + and **dropped `BlogNames` entirely**. There is no compatibility view, so naming it is + `no such table`. Do not recreate it +- **Joining `Notes` to `Blogs`**: `FROM Blogs B INNER JOIN Notes N ON N.NoteBlogId = B.BlogId` +- **Joining `Notes` to `Posts` also goes through `Blogs`**, since `Posts` has only + `BlogName`: `Posts P JOIN Blogs RB ON RB.BlogName = P.BlogName JOIN Notes N ON + N.RootBlogId = RB.BlogId` (`GetRepliesWithFilledText`) +- **Resolve a name by filtering `Blogs`, never by scanning `Notes`**: + `WHERE NoteBlogId = (SELECT BlogId FROM Blogs WHERE BlogName = @name)`. The subquery is a + primary-key probe and does not show against the 1.2M-row table +- **`AddNote` registers both blogs *and* the note type** before inserting, all in one + transaction. `RegisterBlog` does `INSERT OR IGNORE` into `Blogs`, then assigns + `BlogId = MAX(BlogId) + 1` where it is NULL. Unlike `AddBlog`, it does not skip `deact` + names, because a note by a deactivated blog still needs an ID. `NoteTypes` is a table + rather than a `CHECK` constraint precisely so an unseen type is an `INSERT`. Without + that registration a type would resolve to `NULL` and fail the `NOT NULL` on `TypeId`, + losing the note +- **Assigning a `BlogId` is bookkeeping and must not move `DateModified`** +- **`Blogs.BlogId` is NULL on ~166k of ~199k rows**, every blog that has never appeared in + a note. An inner join on it silently drops them. Correct for engagement queries, wrong + for anything listing the registry. `ix_Blogs_BlogId` is `UNIQUE`, which allows many NULLs +- **IDs are stable and must never be renumbered.** They are stored in 1.2M `Notes` rows. + Triggers `trg_Blogs_BlogId_NoDelete` and `trg_Blogs_BlogId_Immutable` abort any + `DELETE` of a `Blogs` row that has a `BlogId`, and any change to its `BlogId` or + `BlogName`. A blog renamed upstream gets a new row. Remove a blog with `IsActive = 0`. + These triggers are also what make `MAX(BlogId) + 1` safe: no ID can ever be freed for + reuse - Prefer `TypeId = (SELECT TypeId FROM NoteTypes WHERE Type = 'reply')` over a hardcoded ID. A negated `TypeId NOT IN (SELECT …)` is only correct because `TypeId` is `NOT NULL` - Duplicate-key detection uses `IsNotesDuplicateKey`, which matches the constraint and the diff --git a/RERUN.sqbpro b/RERUN.sqbpro index 0134308..a05f74e 100644 --- a/RERUN.sqbpro +++ b/RERUN.sqbpro @@ -1,59 +1,69 @@ -
UPDATE Posts -SET HasNotesGathered = 0 -WHERE (BlogName, PostID) IN ( - SELECT p.BlogName, p.PostID - FROM Posts p - WHERE p.HasNotesGathered = 1 - AND P.notesGatheredDatetime < 1774294520 - AND EXISTS ( - SELECT 1 - FROM Notes n - WHERE n.PostID = p.PostID - AND n.RootBlogId = (SELECT BlogId FROM BlogNames WHERE BlogName = p.BlogName) - --AND n.TypeId NOT IN (SELECT TypeId FROM NoteTypes WHERE Type IN ('reblog', 'reply')) - ) - ORDER BY P.PostDate ASC - --LIMIT 500 -);select * +
select * from Blogs --update blogs set HasBeenOutput = 1 where HasBeenOutput = 0 AND blogname in -( -'teaberrybee', -'reddevilgoddesstoo', -'waywardog13', -'wzjustbrowsing-blog', -'lewerta', -'nudenymph', -'caylachief' - +('udontn33dh1m', +'tyrantsxblood', +'sentry-34', +'deathcabforfrankie', +'abheith-sasta', +'kuwaiikittenghost', +'kansasmud', +'03diesel', +'itzameallieee', +'fireball-temptations', +'mamaisamess', +'906raised-and-dogobsessed', +'the-queerist-wolf', +'counting-corpsess', +'aqueenbby', +'maybememoriesx', +'queenofnevers', +'obsidian-psyche', +'lilmissellexo', +'alittlebunny95', +'rage--and--grace', +'savage-deniz', +'daddyspuddleprincess', +'littledefenstration', +'bearded-snorlax', +'thosesummerskiess', +'tubadtoph', +'lieutenant-dan-ice-cream', +'brittvnybitch', +'a-smol-gayologist', +'sum1random', +'samsternelly', +'littlemouseylauren', +'princessleiaorgasma', +'bloodstaineddkisses', +'letsfacerealitybabe', +'x--marks--thespot', +'space-and-suffering', +'rinarootski', +'thiccandtired', +'fvcking-scvmbag', +'fullblownwizard', +'bigjewface', +'unleash-the-krayken', +'bumpintheroad', +'liltexasjedii', +'nawtydude', +'queenpeachqueen', +'the-clansman', +'balmain-bxtch' )select P.slug, N.replyText, rbn.BlogName as RootBlogName, n.PostID, nbn.BlogName || '.tumblr.com' as NoteBlogName, DatetimeCrawled, TimeStamp, nt.Type, rbn.BlogName || '.tumblr.com/post/' || n.postid, datetime(timestamp, 'unixepoch') from Notes N inner join Posts P on p.PostID = n.PostID -inner join BlogNames rbn on rbn.BlogId = n.RootBlogId -inner join BlogNames nbn on nbn.BlogId = n.NoteBlogId +inner join Blogs rbn on rbn.BlogId = n.RootBlogId +inner join Blogs nbn on nbn.BlogId = n.NoteBlogId inner join NoteTypes nt on nt.TypeId = n.TypeId where DatetimeCrawled > '2026-08-07 11:47:22' and nt.Type like 'r%' and P.IsActive = 1 -order by n.DatetimeCrawledSELECT distinct - '''' || blogname || ''',', - blogs.* - , blogname || '.tumblr.com' -FROM - Blogs - inner JOIN - Notes on notes.noteBlogId = blogs.BlogId - inner JOIN - NoteTypes on NoteTypes.TypeId = Notes.TypeId -WHERE - HasBeenOutput = 0 and NoteTypes.Type = 'reblog' -order by - NoteTypes.Type desc, - DateAdded desc -LIMIT 100;WITH ReplyCounts AS ( +order by n.DatetimeCrawledWITH ReplyCounts AS ( SELECT NoteBlogId, COUNT(DISTINCT replyText) AS DistinctReplyCount @@ -68,13 +78,13 @@ SELECT c.DistinctReplyCount FROM Notes n JOIN ReplyCounts c ON n.NoteBlogId = c.NoteBlogId -JOIN BlogNames rbn ON rbn.BlogId = n.RootBlogId -JOIN BlogNames nbn ON nbn.BlogId = n.NoteBlogId +JOIN Blogs rbn ON rbn.BlogId = n.RootBlogId +JOIN Blogs nbn ON nbn.BlogId = n.NoteBlogId JOIN NoteTypes t ON t.TypeId = n.TypeId where replyText <> '.' and t.Type <> 'reply' --AND nbn.BlogName NOT IN ( 'roadblocker21', 'thesaddemon666', 'edwardabbeyhoffman', 'tattedsoldier20', 'zomb-eh', 'animalistic13', 'indken', 'maccloud1592', --'moss-wizard', 'supertrucker12682', 'exploringthrupics', 'padeyepete' ) -order by c.DistinctReplyCount desc, nbn.BlogName, n.DateModified desc, replyText, rbn.BlogName, PostIDWITH PostsWithCount AS ( SELECT P.BlogName, P.PostID, 1925013599 AS LatestNoteTimestamp, P.NotesGatheredDateTime, COUNT(P.PostID) OVER(PARTITION BY P.BlogName) AS CNT, P.HasNotesGathered, P.NotFound, P.PostDate FROM Posts P WHERE COALESCE(P.IsActive, 1) = 1 ), Unioned AS ( SELECT BlogName, PostID, LatestNoteTimestamp, NotesGatheredDateTime, CNT, PostDate FROM PostsWithCount WHERE NotFound = 0 AND HasNotesGathered = 0 UNION SELECT BlogName, PostID, LatestNoteTimestamp, NotesGatheredDateTime, CNT, PostDate FROM PostsWithCount WHERE BlogName = 'zomb-eh' AND NotFound = 0 AND NotesGatheredDateTime < unixepoch('now', 'localtime', '-3 days') ) SELECT U.BlogName, U.PostID, U.LatestNoteTimestamp, U.NotesGatheredDateTime, U.CNT FROM Unioned U WHERE (U.NotesGatheredDateTime < 1786134037 OR U.NotesGatheredDateTime IS NULL) ORDER BY U.NotesGatheredDateTime, U.PostDate DESC, U.BlogName, U.PostID;delete from posts where postid in +order by c.DistinctReplyCount desc, nbn.BlogName, n.DateModified desc, replyText, rbn.BlogName, PostIDWITH PostsWithCount AS ( SELECT P.BlogName, P.PostID, 1925013599 AS LatestNoteTimestamp, P.NotesGatheredDateTime, COUNT(P.PostID) OVER(PARTITION BY P.BlogName) AS CNT, P.HasNotesGathered, P.NotFound, P.PostDate FROM Posts P WHERE COALESCE(P.IsActive, 1) = 1 ), Unioned AS ( SELECT BlogName, PostID, LatestNoteTimestamp, NotesGatheredDateTime, CNT, PostDate FROM PostsWithCount WHERE NotFound = 0 AND HasNotesGathered = 0 UNION SELECT BlogName, PostID, LatestNoteTimestamp, NotesGatheredDateTime, CNT, PostDate FROM PostsWithCount WHERE BlogName = 'zomb-eh' AND NotFound = 0 AND NotesGatheredDateTime < unixepoch('now', 'localtime', '-3 days') ) SELECT U.BlogName, U.PostID, U.LatestNoteTimestamp, U.NotesGatheredDateTime, U.CNT FROM Unioned U WHERE (U.NotesGatheredDateTime < 1787237598 OR U.NotesGatheredDateTime IS NULL) ORDER BY U.NotesGatheredDateTime, U.PostDate DESC, U.BlogName, U.PostID;delete from posts where postid in ( '741662499571728384', 178892849664, @@ -82,28 +92,324 @@ order by c.DistinctReplyCount desc, nbn.BlogName, n.DateModified desc, replyText 177012868749, 169950081964, 755440787056099328 -)select * +)select * -- delete from notes -where postid not in (select distinct postid from posts where IsActive = 1)SELECT - * -FROM - POSTS P -WHERE - P.ByLikes = 1 - AND - P.DateCreated > '2026-05-26 17:47:32' -ORDER BY - P.DateCreated descupdate posts set IsActive = 0 where blogname IN ( 'shoebiedoo', 'redheaded-girlygirl', 'xlittle-ghost' )update Posts -set IsActive = 0 -where postid in -( - - -'731937314675310592' - +where postid not in (select distinct postid from posts where IsActive = 1)-- ============================================================================ +-- blogs-added-after-august-2026-with-reblog-or-reply.sql +-- +-- Purpose: Of all blogs in Blogs, find the ones that show up in Notes as the +-- engager (NoteBlogId) on a 'reblog' or 'reply' note, sorted by +-- DateAdded. (Originally scoped to "added after August 2026" -- +-- that cutoff is now removed per request; QUERY 2 shows how to put +-- a date floor back if needed.) +-- +-- Read-only. No INSERT/UPDATE/DELETE/DDL anywhere in this file. +-- +-- How to use (DB Browser for SQLite): +-- 1. File > Open Database -> TL.db +-- 2. Execute SQL tab. Each QUERY below is independent; Ctrl+Enter runs just +-- the one your cursor is in. +-- +-- The join, once: +-- "Added after August 2026" filters Blogs.DateAdded. "Has a reblog/reply +-- note" means the blog is the engager, which is NoteBlogId -- not +-- RootBlogId, which is the blog that *owns* the post being reacted to +-- (see find-notes-on-inactive-posts.sql for that side). Per TL.db.md, the +-- Blogs<->Notes join is a single integer hop and should not be routed +-- through Blogs: +-- Blogs.BlogId = Notes.NoteBlogId +-- EXISTS is used rather than a JOIN so a blog with many qualifying notes +-- still contributes one output row. +-- +-- Excluding notes on an inactive post: same shape as +-- find-notes-on-inactive-posts.sql -- Notes only carries RootBlogId (an +-- integer), so reaching Posts.IsActive needs the one text hop the rest of +-- this file avoids: RootBlogId -> Blogs.BlogName = Posts.BlogName, +-- matched on PostID. It's a LEFT JOIN, not an inner one: only 3,867 of +-- 20,430 root/engager blogs have any stored Posts rows at all (TL.db.md), +-- so most reblog/reply notes have no Posts row to check and must be kept, +-- not dropped by an inner join. COALESCE(p.IsActive, 1) = 1 keeps a note +-- unless its post is explicitly IsActive = 0 -- NULL (no Posts row, or a +-- stored row with no flag written) means live, per the schema's own +-- convention (TL.db.md, "Posts.IsActive and Notes.IsActive"). This is a +-- big filter in practice: of the blogs that qualified before it, most +-- have every one of their reblog/reply notes pointing at a since-removed +-- post, not just some -- verified against the live data, not assumed. +-- +-- On DateAdded: this column is not written consistently -- most rows hold +-- ISO 'yyyy-MM-dd HH:mm:ss', but 17k+ hold US 'M/d/yy' from a 2025 bulk +-- import (see TL.db.md, "DateAdded is not written consistently"). As text, +-- those two shapes do not sort or compare against each other correctly, so +-- QUERY 0 normalises both to an ISO date before filtering. In the live data +-- every US-format row predates August 2026 anyway (only '12/23/25' and +-- '12/24/25' occur), so this makes no difference to the current answer -- +-- it's here so the query stays correct if that ever changes. +-- ============================================================================ +-- ---------------------------------------------------------------------------- +-- QUERY 0 / THE ANSWER -- one row per qualifying blog, sorted by DateAdded +-- descending (normalised -- see the note above). No date cutoff, but now +-- scoped to HasBeenOutput = 0 AND IsActive = 1. 4,739 rows in the live +-- data. +-- +-- earliest_reblog_or_reply_utc is the MIN(TimeStamp) among this blog's +-- reblog-or-reply notes (either type counts -- see the column name). +-- Getting this meant switching QUERY 0 from EXISTS to an inner JOIN + +-- GROUP BY: EXISTS can only tell you a qualifying row is present, not +-- aggregate over which ones. No CASE is needed inside the MIN() because +-- the WHERE below already restricts the joined rows to reblog/reply, so +-- every row a blog brings into the aggregate is one this column should +-- consider. A blog appears exactly once, same as before, and this column +-- is never NULL for a row that's in the result at all (an earlier +-- revision aggregated reblog only, which left it NULL for the 181 blogs +-- that had replies but no reblogs). +-- ---------------------------------------------------------------------------- +WITH BlogsSplit AS ( + SELECT + b.BlogId, + b.BlogName, + b.DateAdded, + CASE WHEN b.DateAdded LIKE '____-__-__%' THEN 1 ELSE 0 END AS IsIso, + -- for the US 'M/d/yy' shape only: everything after the first '/' + substr(b.DateAdded, instr(b.DateAdded, '/') + 1) AS RestAfterMonth + FROM Blogs b + WHERE b.BlogId IS NOT NULL and HasBeenOutput = 0 and IsActive = 1 -- a blog can only match Notes if it has one +), +BlogsNorm AS ( + SELECT + BlogId, + BlogName, + DateAdded, + CASE + WHEN IsIso = 1 THEN date(DateAdded) + ELSE date( + '20' || substr(RestAfterMonth, instr(RestAfterMonth, '/') + 1) || '-' || + substr('00' || substr(DateAdded, 1, instr(DateAdded, '/') - 1), -2) || '-' || + substr('00' || substr(RestAfterMonth, 1, instr(RestAfterMonth, '/') - 1), -2) + ) + END AS DateAddedNorm + FROM BlogsSplit ) +SELECT + bn.BlogId, + bn.BlogName, + bn.DateAdded, + bn.DateAddedNorm, + datetime(MIN(n.TimeStamp), 'unixepoch') AS earliest_reblog_or_reply_utc + FROM BlogsNorm bn + JOIN Notes n ON n.NoteBlogId = bn.BlogId + JOIN NoteTypes t ON t.TypeId = n.TypeId + JOIN Blogs root_bn ON root_bn.BlogId = n.RootBlogId + LEFT JOIN Posts p ON p.BlogName = root_bn.BlogName + AND p.PostID = n.PostID + WHERE t.Type IN ('reblog')--, 'reply') + AND COALESCE(p.IsActive, 1) = 1 -- exclude notes on a post explicitly marked removed + GROUP BY bn.BlogId, bn.BlogName, bn.DateAdded, bn.DateAddedNorm + ORDER BY bn.DateAddedNorm desc; -
+ +-- ---------------------------------------------------------------------------- +-- QUERY 1 -- same answer, with a per-blog breakdown of which type(s) fired +-- and how many. Useful once QUERY 0 has rows; redundant while it's empty. +-- ---------------------------------------------------------------------------- +-- WITH BlogsSplit AS ( ... ), BlogsNorm AS ( ... ) -- reuse the CTEs above +-- +-- SELECT +-- bn.BlogId, +-- bn.BlogName, +-- bn.DateAddedNorm, +-- SUM(CASE WHEN t.Type = 'reblog' THEN 1 ELSE 0 END) AS reblog_count, +-- SUM(CASE WHEN t.Type = 'reply' THEN 1 ELSE 0 END) AS reply_count +-- FROM BlogsNorm bn +-- JOIN Notes n ON n.NoteBlogId = bn.BlogId +-- JOIN NoteTypes t ON t.TypeId = n.TypeId +-- WHERE t.Type IN ('reblog', 'reply') +-- GROUP BY bn.BlogId, bn.BlogName, bn.DateAddedNorm +-- ORDER BY bn.DateAddedNorm; + + +-- ---------------------------------------------------------------------------- +-- QUERY 2 -- put a date floor back, if wanted later. +-- Same as QUERY 0, with one extra line in the outer WHERE: +-- AND bn.DateAddedNorm >= '2026-09-01' -- or whatever cutoff +-- ---------------------------------------------------------------------------- +-- ============================================================================ +-- blogs-added-after-august-2026-with-reblog-or-reply.sql +-- +-- Purpose: Of all blogs in Blogs, find the ones that show up in Notes as the +-- engager (NoteBlogId) on a 'reblog' or 'reply' note, sorted by +-- DateAdded. (Originally scoped to "added after August 2026" -- +-- that cutoff is now removed per request; QUERY 2 shows how to put +-- a date floor back if needed.) +-- +-- Read-only. No INSERT/UPDATE/DELETE/DDL anywhere in this file. +-- +-- How to use (DB Browser for SQLite): +-- 1. File > Open Database -> TL.db +-- 2. Execute SQL tab. Each QUERY below is independent; Ctrl+Enter runs just +-- the one your cursor is in. +-- +-- The join, once: +-- "Added after August 2026" filters Blogs.DateAdded. "Has a reblog/reply +-- note" means the blog is the engager, which is NoteBlogId -- not +-- RootBlogId, which is the blog that *owns* the post being reacted to +-- (see find-notes-on-inactive-posts.sql for that side). Per TL.db.md, the +-- Blogs<->Notes join is a single integer hop and should not be routed +-- through Blogs: +-- Blogs.BlogId = Notes.NoteBlogId +-- EXISTS is used rather than a JOIN so a blog with many qualifying notes +-- still contributes one output row. +-- +-- Excluding notes on an inactive post: same shape as +-- find-notes-on-inactive-posts.sql -- Notes only carries RootBlogId (an +-- integer), so reaching Posts.IsActive needs the one text hop the rest of +-- this file avoids: RootBlogId -> Blogs.BlogName = Posts.BlogName, +-- matched on PostID. It's a LEFT JOIN, not an inner one: only 3,867 of +-- 20,430 root/engager blogs have any stored Posts rows at all (TL.db.md), +-- so most reblog/reply notes have no Posts row to check and must be kept, +-- not dropped by an inner join. COALESCE(p.IsActive, 1) = 1 keeps a note +-- unless its post is explicitly IsActive = 0 -- NULL (no Posts row, or a +-- stored row with no flag written) means live, per the schema's own +-- convention (TL.db.md, "Posts.IsActive and Notes.IsActive"). This is a +-- big filter in practice: of the blogs that qualified before it, most +-- have every one of their reblog/reply notes pointing at a since-removed +-- post, not just some -- verified against the live data, not assumed. +-- +-- On DateAdded: this column is not written consistently -- most rows hold +-- ISO 'yyyy-MM-dd HH:mm:ss', but 17k+ hold US 'M/d/yy' from a 2025 bulk +-- import (see TL.db.md, "DateAdded is not written consistently"). As text, +-- those two shapes do not sort or compare against each other correctly, so +-- QUERY 0 normalises both to an ISO date before filtering. In the live data +-- every US-format row predates August 2026 anyway (only '12/23/25' and +-- '12/24/25' occur), so this makes no difference to the current answer -- +-- it's here so the query stays correct if that ever changes. +-- ============================================================================ + + +-- ---------------------------------------------------------------------------- +-- QUERY 0 / THE ANSWER -- one row per qualifying blog, sorted by +-- earliest_reblog_or_reply_utc then DateAdded descending (normalised -- +-- see the note above). No date cutoff, but scoped to HasBeenOutput = 0 +-- AND IsActive = 1, and now excluding notes on a removed post (see the +-- header note above). 1,804 rows in the live data as of this revision -- +-- down from 4,396 just before this exclusion was added, because most of +-- the blogs that dropped out had *every* reblog/reply note pointing at a +-- now-inactive post, not just some (the number moves between runs +-- regardless -- crawling and output flip HasBeenOutput/IsActive on live +-- rows). +-- +-- earliest_reblog_or_reply_utc is the earliest TimeStamp among this +-- blog's reblog-or-reply notes (either type counts -- see the column +-- name); earliest_reblog_or_reply_postid and _root_blogid identify that +-- specific note's post: PostID + RootBlogId together, not PostID alone -- +-- see TL.db.md ("345 post IDs exist under more than one blog"), same +-- caution as in find-notes-on-inactive-posts.sql. Resolve RootBlogId to a +-- name via Blogs (or Blogs, tolerating a miss) if you need it. +-- +-- Getting "which note" rather than just "when" doesn't fit a plain +-- MIN()/GROUP BY -- an aggregate can tell you the earliest value but not +-- which row it came from. EarliestNote instead ranks each blog's +-- reblog/reply notes with ROW_NUMBER() OVER (PARTITION BY NoteBlogId +-- ORDER BY TimeStamp), and QUERY 0 takes rn = 1. The ORDER BY carries a +-- PostID tiebreak because (NoteBlogId, TimeStamp) is not unique in this +-- data -- ties exist (e.g. NoteBlogId 12 has 7 notes at the same +-- TimeStamp) -- so without a tiebreak the "earliest" pick would be +-- arbitrary among ties rather than deterministic. +-- +-- EarliestNote also excludes notes on an inactive post before ranking +-- (see the header note above), so "earliest" means earliest surviving +-- note, not earliest overall -- a blog whose true-earliest note pointed +-- at a since-removed post now surfaces its next-earliest live one +-- instead. Applying the exclusion here, not as a filter on QUERY 0's +-- final rows, matters: filtering after ROW_NUMBER would have picked the +-- removed-post note as rn = 1 and then dropped the whole row instead of +-- promoting the next candidate. +-- ---------------------------------------------------------------------------- +WITH BlogsSplit AS ( + SELECT + b.BlogId, + b.BlogName, + b.DateAdded, + CASE WHEN b.DateAdded LIKE '____-__-__%' THEN 1 ELSE 0 END AS IsIso, + -- for the US 'M/d/yy' shape only: everything after the first '/' + substr(b.DateAdded, instr(b.DateAdded, '/') + 1) AS RestAfterMonth + FROM Blogs b + WHERE b.BlogId IS NOT NULL and HasBeenOutput = 0 and IsActive = 1 -- a blog can only match Notes if it has one +), +BlogsNorm AS ( + SELECT + BlogId, + BlogName, + DateAdded, + CASE + WHEN IsIso = 1 THEN date(DateAdded) + ELSE date( + '20' || substr(RestAfterMonth, instr(RestAfterMonth, '/') + 1) || '-' || + substr('00' || substr(DateAdded, 1, instr(DateAdded, '/') - 1), -2) || '-' || + substr('00' || substr(RestAfterMonth, 1, instr(RestAfterMonth, '/') - 1), -2) + ) + END AS DateAddedNorm + FROM BlogsSplit +), +EarliestNote AS ( + SELECT + n.NoteBlogId, + n.RootBlogId, + n.PostID, + n.TimeStamp, + ROW_NUMBER() OVER ( + PARTITION BY n.NoteBlogId + ORDER BY n.TimeStamp ASC, n.PostID ASC + ) AS rn + FROM Notes n + JOIN NoteTypes t ON t.TypeId = n.TypeId + JOIN Blogs root_bn ON root_bn.BlogId = n.RootBlogId + LEFT JOIN Posts p ON p.BlogName = root_bn.BlogName + AND p.PostID = n.PostID + WHERE t.Type IN ('reblog')--, 'reply') + AND n.NoteBlogId IN (SELECT BlogId FROM BlogsNorm) -- scope the window to blogs we care about + AND COALESCE(p.IsActive, 1) = 1 -- exclude notes on a post explicitly marked removed +) +SELECT + bn.BlogId, + bn.BlogName || '.tumblr.com', + '''' || bn.blogname || ''',', + bn.DateAdded, + bn.DateAddedNorm, + datetime(en.TimeStamp, 'unixepoch') AS earliest_reblog_or_reply_utc, + en.PostID AS earliest_reblog_or_reply_postid, + en.RootBlogId AS earliest_reblog_or_reply_root_blogid + FROM BlogsNorm bn + JOIN EarliestNote en ON en.NoteBlogId = bn.BlogId AND en.rn = 1 + ORDER BY earliest_reblog_or_reply_utc, bn.DateAddedNorm desc + limit 50; + + +-- ---------------------------------------------------------------------------- +-- QUERY 1 -- same answer, with a per-blog breakdown of which type(s) fired +-- and how many. Useful once QUERY 0 has rows; redundant while it's empty. +-- ---------------------------------------------------------------------------- +-- WITH BlogsSplit AS ( ... ), BlogsNorm AS ( ... ) -- reuse the CTEs above +-- +-- SELECT +-- bn.BlogId, +-- bn.BlogName, +-- bn.DateAddedNorm, +-- SUM(CASE WHEN t.Type = 'reblog' THEN 1 ELSE 0 END) AS reblog_count, +-- SUM(CASE WHEN t.Type = 'reply' THEN 1 ELSE 0 END) AS reply_count +-- FROM BlogsNorm bn +-- JOIN Notes n ON n.NoteBlogId = bn.BlogId +-- JOIN NoteTypes t ON t.TypeId = n.TypeId +-- WHERE t.Type IN ('reblog', 'reply') +-- GROUP BY bn.BlogId, bn.BlogName, bn.DateAddedNorm +-- ORDER BY bn.DateAddedNorm; + + +-- ---------------------------------------------------------------------------- +-- QUERY 2 -- put a date floor back, if wanted later. +-- Same as QUERY 0, with one extra line in the outer WHERE: +-- AND bn.DateAddedNorm >= '2026-09-01' -- or whatever cutoff +-- ---------------------------------------------------------------------------- +
diff --git a/TL.sqbpro b/TL.sqbpro index 8d0d93d..94606e1 100644 --- a/TL.sqbpro +++ b/TL.sqbpro @@ -15,9 +15,9 @@ ORDER BY FROM Notes N inner JOIN - BlogNames rbn on rbn.BlogId = N.RootBlogId + Blogs rbn on rbn.BlogId = N.RootBlogId inner JOIN - BlogNames nbn on nbn.BlogId = N.NoteBlogId + Blogs nbn on nbn.BlogId = N.NoteBlogId inner JOIN NoteTypes t on t.TypeId = N.TypeId inner JOIN diff --git a/URLNotesGrabberCORE/DataAccess.cs b/URLNotesGrabberCORE/DataAccess.cs index 1924966..ad4122a 100644 --- a/URLNotesGrabberCORE/DataAccess.cs +++ b/URLNotesGrabberCORE/DataAccess.cs @@ -216,18 +216,19 @@ namespace URLNotesGrabberCORE #region Notes integer schema // Notes stopped storing names on 2026-08-07: RootBlogName/NoteBlogName/Type became - // RootBlogId/NoteBlogId/TypeId, resolved through BlogNames and NoteTypes. There is no + // RootBlogId/NoteBlogId/TypeId, resolved through Blogs.BlogId and NoteTypes. There is no // compatibility view -- a query naming an old column fails outright, so this is a hard // cut rather than an optional column like IsActive. See TL.db.md. // + // Blogs.BlogId is the only ID authority as of 2026-09-28; the BlogNames table that used + // to hold the IDs is gone. Every name in Notes has a Blogs row, created by RegisterBlog. + // // Two shapes recur below and are spelled out inline rather than hidden behind a helper, // so that every statement reads as the SQL it actually runs: - // (SELECT BlogId FROM BlogNames WHERE BlogName = @name) -- unique-index probe, 20k rows + // (SELECT BlogId FROM Blogs WHERE BlogName = @name) -- primary-key probe // (SELECT TypeId FROM NoteTypes WHERE Type = 'reply') -- 5 rows, effectively free - // Joining Notes to Blogs is the one case that must NOT route through BlogNames: Blogs - // carries its own BlogId, so N.NoteBlogId = B.BlogId is a single integer hop. Joining - // Notes to Posts is the opposite case -- Posts has only BlogName, so it has to go - // through BlogNames. + // Joining Notes to Blogs is N.NoteBlogId = B.BlogId, a single hop on the unique index. + // Joining Notes to Posts goes through Blogs too -- Posts has only BlogName. /// /// True when the exception is a duplicate-key collision on Notes. The message embeds the @@ -242,14 +243,30 @@ namespace URLNotesGrabberCORE } /// - /// Gives a blog name an ID if it does not have one. No read-back and no round trip -- a - /// name that is already registered keeps the ID that 1.18M Notes rows point at. + /// Ensures a blog has a Blogs row and a BlogId, so a note can point at it. No read-back + /// and no round trip -- a blog that already has an ID keeps the one Notes rows point at. + /// Unlike AddBlog this does not skip "deact" names: a note by a deactivated blog still + /// needs an ID, and Blogs is the only place one can live. + /// Assigning the ID is bookkeeping, not a content change, so DateModified is not touched. + /// MAX(BlogId) + 1 cannot hand out a used ID because trg_Blogs_BlogId_NoDelete stops any + /// row that holds one from being deleted. /// - private static void RegisterBlogName(SQLiteConnection connection, SQLiteTransaction? transaction, string blogName) + private static void RegisterBlog(SQLiteConnection connection, SQLiteTransaction? transaction, string blogName) { - using SQLiteCommand command = new SQLiteCommand("INSERT OR IGNORE INTO BlogNames (BlogName) VALUES (@BlogName)", connection, transaction); - command.Parameters.AddWithValue("@BlogName", blogName); - command.ExecuteNonQuery(); + string now = DateTime.Now.ToString("yyyy-MM-dd HH:mm:ss"); + + using (SQLiteCommand command = new SQLiteCommand("INSERT OR IGNORE INTO Blogs (BlogName, DateAdded, DateModified, DateCreated) VALUES (@BlogName, @Now, @Now, @Now)", connection, transaction)) + { + command.Parameters.AddWithValue("@BlogName", blogName); + command.Parameters.AddWithValue("@Now", now); + command.ExecuteNonQuery(); + } + + using (SQLiteCommand command = new SQLiteCommand("UPDATE Blogs SET BlogId = (SELECT IFNULL(MAX(BlogId), 0) + 1 FROM Blogs) WHERE BlogName = @BlogName AND BlogId IS NULL", connection, transaction)) + { + command.Parameters.AddWithValue("@BlogName", blogName); + command.ExecuteNonQuery(); + } } /// @@ -762,7 +779,6 @@ namespace URLNotesGrabberCORE { DBPath ??= GetDefaultDbPath(); //try { AddPost(rootBlogName, postID, DBPath); } catch { } - try { AddBlog(noteBlogName, false, DBPath); } catch { } using SQLiteConnection connection2 = new SQLiteConnection("Data Source=" + DBPath); int rowsInserted = 0; @@ -771,23 +787,24 @@ namespace URLNotesGrabberCORE { connection2.Open(); - // Notes stores integer IDs, so both participants and the type have to exist in - // their lookup table before the note can point at them. + // Notes stores integer IDs, so both participants need a Blogs row with a BlogId, + // and the type a NoteTypes row, before the note can point at them. RegisterBlog + // also does what the AddBlog call here used to: register the note's blog. // - // All four statements run in one transaction so a crash cannot leave a name or a + // All statements run in one transaction so a crash cannot leave a blog or a // type registered with no note. The transaction is committed before the console // output below, which sleeps -- a write lock must not be held across that. using (SQLiteTransaction transaction = connection2.BeginTransaction()) { - RegisterBlogName(connection2, transaction, rootBlogName); - RegisterBlogName(connection2, transaction, noteBlogName); + RegisterBlog(connection2, transaction, rootBlogName); + RegisterBlog(connection2, transaction, noteBlogName); RegisterNoteType(connection2, transaction, type ?? string.Empty); // INSERT OR IGNORE, and no IsActive in the column list: re-crawling a note // that was removed elsewhere leaves the existing row -- and its flag -- alone. string sql = "INSERT OR IGNORE INTO Notes (RootBlogId, NoteBlogId, PostID, TimeStamp, TypeId, DatetimeCrawled, DateModified, DateCreated) " + - "SELECT (SELECT BlogId FROM BlogNames WHERE BlogName = @rootBlogName), " + - " (SELECT BlogId FROM BlogNames WHERE BlogName = @noteBlogName), " + + "SELECT (SELECT BlogId FROM Blogs WHERE BlogName = @rootBlogName), " + + " (SELECT BlogId FROM Blogs WHERE BlogName = @noteBlogName), " + " @PostID, @TimeStamp, " + " (SELECT TypeId FROM NoteTypes WHERE Type = @Type), " + " @DatetimeCrawled, @DateModified, @DateCreated"; @@ -1126,11 +1143,11 @@ namespace URLNotesGrabberCORE { connection.Open(); - string sql = "SELECT DISTINCT BN.BlogName as blogName, N.PostID" + + string sql = "SELECT DISTINCT RB.BlogName as blogName, N.PostID" + " FROM Notes N" + - " INNER JOIN BlogNames BN ON BN.BlogId = N.RootBlogId" + + " INNER JOIN Blogs RB ON RB.BlogId = N.RootBlogId" + " WHERE N.TypeId = (SELECT TypeId FROM NoteTypes WHERE Type = 'reply')" + AndIsActive("Notes", "N", DBPath) + - " ORDER BY BN.BlogName, N.PostID"; + " ORDER BY RB.BlogName, N.PostID"; using (SQLiteCommand command = new SQLiteCommand(sql, connection)) { @@ -1171,11 +1188,11 @@ namespace URLNotesGrabberCORE connection.Open(); // Grouped on the integer rather than the name: the group key is what gets sorted, - // and BN.BlogName comes along for free off the join. - string sql = @"SELECT BN.BlogName as blogName, N.PostID, + // and RB.BlogName comes along for free off the join. + string sql = @"SELECT RB.BlogName as blogName, N.PostID, MAX(N.TimeStamp) as LatestTimestamp FROM Notes N - INNER JOIN BlogNames BN ON BN.BlogId = N.RootBlogId + INNER JOIN Blogs RB ON RB.BlogId = N.RootBlogId WHERE N.TypeId = (SELECT TypeId FROM NoteTypes WHERE Type = 'reply') AND (N.replyText IS NULL OR N.replyText = '' OR N.replyText = '.')" + AndIsActive("Notes", "N", DBPath) + @" GROUP BY N.RootBlogId, N.PostID @@ -1220,13 +1237,13 @@ namespace URLNotesGrabberCORE { connection.Open(); - // Posts carries only BlogName, so this is the one join to Notes that has to go - // through BlogNames -- there is no Posts.BlogId to hop on. The name predicate is - // pushed into the 20k-row lookup, which then feeds integers to the Notes key. + // Posts carries only BlogName, so the join to Notes goes through Blogs -- there is + // no Posts.BlogId to hop on. Each post's name is a primary-key probe on Blogs, + // which then feeds an integer to the Notes key. string sql = @"SELECT DISTINCT P.BlogName, P.PostID, MAX(N.TimeStamp) as LatestTimestamp FROM Posts P - INNER JOIN BlogNames RBN ON RBN.BlogName = P.BlogName - INNER JOIN Notes N ON N.RootBlogId = RBN.BlogId AND N.PostID = P.PostID + INNER JOIN Blogs RB ON RB.BlogName = P.BlogName + INNER JOIN Notes N ON N.RootBlogId = RB.BlogId AND N.PostID = P.PostID WHERE P.NotFound = 0 AND N.TypeId = (SELECT TypeId FROM NoteTypes WHERE Type = 'reply') AND (N.replyText IS NULL OR N.replyText = '' OR N.replyText = '.')" + AndIsActive("Posts", "P", DBPath) + AndIsActive("Notes", "N", DBPath) + @" @@ -1415,8 +1432,7 @@ namespace URLNotesGrabberCORE try { connection.Open(); - // Blogs is reached in one integer hop off Blogs.BlogId, not through BlogNames -- - // that would add a hop and end in the text comparison the migration removed. + // Blogs is reached in one integer hop off Blogs.BlogId. // The negated form is only correct because Notes.TypeId is NOT NULL. string sql = ""; if (reblogsOnly) @@ -1703,8 +1719,8 @@ namespace URLNotesGrabberCORE connection.Open(); string sql = "UPDATE Notes SET TimeStamp = @timestamp, DateModified = @dateModified " + - "WHERE RootBlogId = (SELECT BlogId FROM BlogNames WHERE BlogName = @rootBlogName) " + - "AND NoteBlogId = (SELECT BlogId FROM BlogNames WHERE BlogName = @noteBlogName) " + + "WHERE RootBlogId = (SELECT BlogId FROM Blogs WHERE BlogName = @rootBlogName) " + + "AND NoteBlogId = (SELECT BlogId FROM Blogs WHERE BlogName = @noteBlogName) " + "AND PostID = @postID AND IFNULL(TimeStamp, 0) <> @timestamp"; using (SQLiteCommand command = new SQLiteCommand(sql, connection)) { @@ -2018,7 +2034,7 @@ namespace URLNotesGrabberCORE // Only fan out to rows that match the SELECT criteria in GetRepliesWithFilledText (NULL/empty/legacy-'.'). Never overwrite '?' (confirmed-empty) or already-fetched text. // The ABS() term cannot use an index on TimeStamp, before or after the integer schema; the NoteBlogId probe is what keeps this off a full scan. string sql = "UPDATE Notes SET replyText = @replyText, DateModified = @dateModified " + - "WHERE NoteBlogId = (SELECT BlogId FROM BlogNames WHERE BlogName = @noteBlogName) " + + "WHERE NoteBlogId = (SELECT BlogId FROM Blogs WHERE BlogName = @noteBlogName) " + "AND ABS(TimeStamp - @TimeStamp) <= 5 " + "AND TypeId = (SELECT TypeId FROM NoteTypes WHERE Type = 'reply') " + "AND (replyText IS NULL OR replyText = '' OR replyText = '.') " + @@ -2064,7 +2080,7 @@ namespace URLNotesGrabberCORE connection.Open(); string sql = "UPDATE Notes SET replyText = @replyText, DateModified = @dateModified " + - "WHERE RootBlogId = (SELECT BlogId FROM BlogNames WHERE BlogName = @rootBlogName) " + + "WHERE RootBlogId = (SELECT BlogId FROM Blogs WHERE BlogName = @rootBlogName) " + "AND PostID = @PostID " + "AND TypeId = (SELECT TypeId FROM NoteTypes WHERE Type = 'reply') " + "AND IFNULL(replyText, '.') <> @replyText"; diff --git a/URLNotesGrabberCORE/TL.db.md b/URLNotesGrabberCORE/TL.db.md index cc20372..ba695b4 100644 --- a/URLNotesGrabberCORE/TL.db.md +++ b/URLNotesGrabberCORE/TL.db.md @@ -24,21 +24,43 @@ Everything below was read out of the live file, not inferred from code. Counts a > Applied by `../normalize-notes.sql`, which took the file from 207 MB to 148 MB. An > earlier change the same day (`../shrink-db.sql`) took it from 267 MB to 207 MB. +> ### ⚠ Breaking change, 2026-09-28: `BlogNames` is gone; `Blogs.BlogId` is the only ID authority +> +> The IDs in `Notes` used to live in a `BlogNames` table, with a copy in `Blogs.BlogId`. +> Nothing kept the copy current, so by 2026-09-28 12,238 blogs first seen after the +> migration had `Blogs.BlogId = NULL`. Every `Notes`-to-`Blogs` join on `BlogId` silently +> skipped them and their 23,148 notes, which kept them out of `GetBlogs`. +> +> `../retire-blognames.sql` fixed this by giving every note participant a `Blogs` row, +> backfilling the IDs (none renumbered), making `ix_Blogs_BlogId` unique, and **dropping +> `BlogNames`**. There is no compatibility view: any query naming it fails with +> `no such table: BlogNames`. Two triggers now protect the IDs. +> +> **Porting an app:** replace `BlogNames` with `Blogs` everywhere. The columns you used, +> `BlogId` and `BlogName`, exist there with the same meaning. A name lookup +> (`SELECT BlogId FROM Blogs WHERE BlogName = ?`) is a primary-key probe, and an ID +> lookup or join (`JOIN Blogs b ON b.BlogId = n.NoteBlogId`) uses the unique +> `ix_Blogs_BlogId`. Every ID in `Notes` resolves to exactly one `Blogs` row. `Blogs.BlogId` +> is **no longer** a stale copy, so any code or docs that distrust it can drop that +> caveat. Never write `BlogId` or `BlogName` on a row that has an ID, and never delete such +> a row: the triggers reject all three. See [`Blogs`](#blogs). + --- ## The three content tables | Table | Rows | What it is | |---|--:|---| -| `Blogs` | 188,620 | The crawl registry — one row per known blog, plus crawl-state flags | +| `Blogs` | 198,560 | The crawl registry, one row per known blog plus crawl-state flags. Also the ID authority for blogs in `Notes` | | `Posts` | 22,468 | Stored post content. Only 3,867 blogs actually have any | -| `Notes` | 1,182,333 | The engagement graph: `NoteBlogId` acted on `(RootBlogId, PostID)` | +| `Notes` | 1,234,830 | The engagement graph: `NoteBlogId` acted on `(RootBlogId, PostID)` | -…supported by two lookup tables that exist only to keep `Notes` small: +(`Blogs` and `Notes` counts as of 2026-09-28; the rest as of 2026-08-07.) + +…supported by one lookup table that exists only to keep `Notes` small: | Table | Rows | What it is | |---|--:|---| -| `BlogNames` | 20,430 | `BlogId` ⇄ `BlogName`. The ID authority for everything in `Notes` | | `NoteTypes` | 5 | `TypeId` ⇄ `Type`. `like`, `reblog`, `reply`, `posted`, `post_attribution` | The engagement graph is the interesting part. 20,311 distinct blogs appear as engagers — @@ -66,27 +88,47 @@ CREATE TABLE "Blogs" ( PRIMARY KEY("BlogName") ); -CREATE INDEX ix_Blogs_BlogId ON Blogs (BlogId); +CREATE UNIQUE INDEX ix_Blogs_BlogId ON Blogs (BlogId); + +CREATE TRIGGER trg_Blogs_BlogId_NoDelete -- no DELETE of a row that has a BlogId +CREATE TRIGGER trg_Blogs_BlogId_Immutable -- no change to its BlogId or BlogName ``` `BlogName` is the primary key, so it is the only indexed way in by name. There is no index -on any flag or date — filtering or sorting on those scans all 188k rows, which is +on any flag or date. Filtering or sorting on those scans the whole table, which is affordable here and is not on `Notes`. -**`BlogId` is new as of 2026-08-07 and is the join key to `Notes`.** It exists so that -`Notes` can reach `Blogs` in a single integer hop rather than going through `BlogNames` -and ending in a text comparison: +**`BlogId` is the ID that `Notes.RootBlogId` and `Notes.NoteBlogId` store, and `Blogs` is +the only place it lives** (since 2026-09-28; see the banner at the top). The join to +`Notes` is one integer hop on the unique index: ```sql --- what you want FROM Blogs B JOIN Notes N ON N.NoteBlogId = B.BlogId - --- not this -FROM Blogs B JOIN BlogNames BN ON BN.BlogName = B.BlogName - JOIN Notes N ON N.NoteBlogId = BN.BlogId ``` -**`BlogId` is NULL on 168,202 of 188,620 rows** — every blog that has never appeared in a +**Every blog that appears in `Notes` has a `Blogs` row with a `BlogId`.** `AddNote` +guarantees it through `RegisterBlog`, which runs in the note's own transaction: + +```sql +INSERT OR IGNORE INTO Blogs (BlogName, DateAdded, DateModified, DateCreated) +VALUES (@name, @now, @now, @now); +UPDATE Blogs SET BlogId = (SELECT IFNULL(MAX(BlogId), 0) + 1 FROM Blogs) + WHERE BlogName = @name AND BlogId IS NULL; +``` + +- Unlike `AddBlog`, this does **not** skip names containing `deact`. A note by a + deactivated blog still needs an ID, so such blogs now get registry rows too, with the + usual defaults (`HasBeenOutput = 0`, `IsActive` left at its default). +- Assigning a `BlogId` is bookkeeping, so it **does not move `DateModified`**. +- `MAX(BlogId) + 1` is safe only because an ID can never be freed. The two triggers see + to that: deleting a row that has a `BlogId`, or changing its `BlogId` or `BlogName`, + aborts. Remove a blog with `IsActive = 0` instead. A blog renamed upstream gets a new + row. Rows with no `BlogId` can still be deleted or renamed freely. +- `INSERT OR REPLACE` on `Blogs` gets around the delete trigger (SQLite does not fire + delete triggers for REPLACE unless `recursive_triggers` is on), and it would wipe the + `BlogId`. It was already forbidden because it resets `IsActive`. Do not use it. + +**`BlogId` is NULL on 165,887 of 198,560 rows**, every blog that has never appeared in a note. That is the large majority, and it is not an error: the registry is far bigger than the engagement graph. An inner join on `BlogId` therefore silently drops those blogs, which is usually what you want for engagement queries and is wrong for registry listings. @@ -187,8 +229,7 @@ CREATE INDEX ix_Notes_NoteBlogId ON Notes (NoteBlogId); **Integer IDs since 2026-08-07 — this is the breaking change.** `RootBlogName`, `NoteBlogName` and `Type` are gone, replaced by `RootBlogId`, `NoteBlogId` and `TypeId`. -Resolve them through [`BlogNames`](#blognames) and [`NoteTypes`](#notetypes), or join -straight to `Blogs` on `BlogId`. The old names were text repeated on 1.18 million rows, +Resolve blog IDs through `Blogs.BlogId` and types through [`NoteTypes`](#notetypes). The old names were text repeated on 1.18 million rows, in the table *and* in every index over it; the swap took the file from 207 MB to 148 MB. The **primary key column order is deliberately unchanged**, so the leading-prefix access @@ -251,43 +292,21 @@ At 1.18M rows this is the table that dictates how the whole database has to be q - `replyText` is `'.'` on 1,167,464 rows — only `reply` notes carry real text. Those dots are inherited from the old column default; new rows get `NULL` instead. -**Resolve IDs by filtering the lookup, not by scanning `Notes`.** The lookup tables are -tiny and uniquely indexed, so pushing a name predicate into them costs nothing and lets -the `Notes` index do the work: +**Resolve IDs by filtering `Blogs`, not by scanning `Notes`.** A name predicate on `Blogs` +is a primary-key probe, so pushing it there costs nothing and lets the `Notes` index do +the work: ```sql --- good: BlogNames resolves the name, then the index is searched +-- good: Blogs resolves the name, then the index is searched SELECT * FROM Notes - WHERE NoteBlogId = (SELECT BlogId FROM BlogNames WHERE BlogName = ?); + WHERE NoteBlogId = (SELECT BlogId FROM Blogs WHERE BlogName = ?); -- also good, same plan SELECT n.* FROM Notes n - JOIN BlogNames b ON b.BlogId = n.NoteBlogId + JOIN Blogs b ON b.BlogId = n.NoteBlogId WHERE b.BlogName = ?; ``` -### `BlogNames` - -```sql -CREATE TABLE BlogNames ( - BlogId INTEGER PRIMARY KEY, - BlogName TEXT NOT NULL UNIQUE -); -``` - -20,430 rows — every name appearing in `Notes` as either participant, and nothing else. -This is the **ID authority**: `Notes.RootBlogId` and `Notes.NoteBlogId` both point here, -and `Blogs.BlogId` is a copy of the value for the blogs that have one. - -**12 of these names have no `Blogs` row.** The registry has never been a superset of the -engagement graph and still is not, so resolving an ID through `Blogs` rather than -`BlogNames` will occasionally find nothing. Use `BlogNames` when you need the name itself -and `Blogs` when you need registry columns. - -IDs are assigned by SQLite and are **stable**: they are stored in over a million `Notes` -rows. Never renumber them. A blog that is renamed upstream should get a new row, not an -edit to an existing one, unless every `Notes` reference is migrated with it. - ### `NoteTypes` ```sql @@ -317,6 +336,9 @@ code to this table's contents, so prefer the join in anything long-lived. ## Porting to the integer schema +> Written for the 2026-08-07 change, and updated for 2026-09-28: wherever this section +> once said `BlogNames`, it now says `Blogs`. `BlogNames` no longer exists. + Everything here was checked against the live 148 MB file. There were 14 affected call sites in `DataAccess.cs` and 16 in `RolodexRepository.cs`. TumblThree needs no changes — its single statement touches `Blogs.IsActive` and `BlogName` only. @@ -333,8 +355,8 @@ the result. | Was | Is now | Resolve via | |---|---|---| -| `Notes.RootBlogName` | `Notes.RootBlogId` | `BlogNames.BlogId` → `.BlogName` | -| `Notes.NoteBlogName` | `Notes.NoteBlogId` | `BlogNames.BlogId` → `.BlogName` | +| `Notes.RootBlogName` | `Notes.RootBlogId` | `Blogs.BlogId` → `.BlogName` | +| `Notes.NoteBlogName` | `Notes.NoteBlogId` | `Blogs.BlogId` → `.BlogName` | | `Notes.Type` | `Notes.TypeId` | `NoteTypes.TypeId` → `.Type` | | `ix_NoteBlogName01` | `ix_Notes_NoteBlogId` | — | @@ -347,14 +369,14 @@ the result. -- was WHERE NoteBlogName = @Name --- now, either form; both search ix_Notes_NoteBlogId after a unique-index lookup -WHERE NoteBlogId = (SELECT BlogId FROM BlogNames WHERE BlogName = @Name) +-- now, either form; both search ix_Notes_NoteBlogId after a primary-key lookup +WHERE NoteBlogId = (SELECT BlogId FROM Blogs WHERE BlogName = @Name) -- or -JOIN BlogNames b ON b.BlogId = n.NoteBlogId WHERE b.BlogName = @Name +JOIN Blogs b ON b.BlogId = n.NoteBlogId WHERE b.BlogName = @Name ``` -Measured 73 ms against 63 ms for the old text form on the busiest blog — the extra hop is -a unique-index probe on a 20k-row table and does not show. +Measured at 73 ms against 63 ms for the old text form on the busiest blog, when the lookup +was still `BlogNames`. The extra hop is one index probe and does not show. ### Joining `Notes` to `Blogs` @@ -364,12 +386,17 @@ This is the join to get right; it is the most common shape in both applications. -- was FROM Blogs B INNER JOIN Notes N ON N.NoteBlogName = B.BlogName --- now: one integer hop, using the new Blogs.BlogId +-- now: one integer hop, using Blogs.BlogId FROM Blogs B INNER JOIN Notes N ON N.NoteBlogId = B.BlogId ``` -Do **not** route this through `BlogNames` — that adds a hop and ends in the text -comparison the change was meant to remove. +Joining `Notes` to `Posts` also goes through `Blogs`, since `Posts` has only a name: + +```sql +FROM Posts P +JOIN Blogs RB ON RB.BlogName = P.BlogName +JOIN Notes N ON N.RootBlogId = RB.BlogId AND N.PostID = P.PostID +``` ### Selecting a name back out @@ -378,12 +405,12 @@ comparison the change was meant to remove. SELECT NoteBlogName AS blogName, COUNT(*) FROM Notes ... GROUP BY NoteBlogName -- now -SELECT bn.BlogName AS blogName, COUNT(*) - FROM Notes n JOIN BlogNames bn ON bn.BlogId = n.NoteBlogId - ... GROUP BY bn.BlogName +SELECT b.BlogName AS blogName, COUNT(*) + FROM Notes n JOIN Blogs b ON b.BlogId = n.NoteBlogId + ... GROUP BY b.BlogName ``` -Group by `n.NoteBlogId` instead of `bn.BlogName` when you only need the name for display — +Group by `n.NoteBlogId` instead of `b.BlogName` when you only need the name for display — grouping on the integer is cheaper and the name comes along for free. ### Filtering by type @@ -407,27 +434,25 @@ because `TypeId` is `NOT NULL`. ### Inserting a note -The crawler must ensure both names have IDs first. `INSERT OR IGNORE` on `BlogNames` is -the whole of it — no read-back, no round trip, safe to run every time: +The crawler must ensure both blogs have IDs first: run the `RegisterBlog` pair shown under +[`Blogs`](#blogs) for each name. No read-back, no round trip, and safe to run every time. +Then: ```sql -INSERT OR IGNORE INTO BlogNames (BlogName) VALUES (@rootBlogName); -INSERT OR IGNORE INTO BlogNames (BlogName) VALUES (@noteBlogName); - INSERT OR IGNORE INTO Notes (RootBlogId, PostID, NoteBlogId, TimeStamp, TypeId, DatetimeCrawled, DateModified, DateCreated) -SELECT (SELECT BlogId FROM BlogNames WHERE BlogName = @rootBlogName), +SELECT (SELECT BlogId FROM Blogs WHERE BlogName = @rootBlogName), @PostID, - (SELECT BlogId FROM BlogNames WHERE BlogName = @noteBlogName), + (SELECT BlogId FROM Blogs WHERE BlogName = @noteBlogName), @TimeStamp, (SELECT TypeId FROM NoteTypes WHERE Type = @Type), @DatetimeCrawled, @DateModified, @DateCreated; ``` Verified: a genuinely new note inserts, and re-running the identical statement inserts 0. -Run all three statements in one transaction so a crash cannot leave a name registered -with no note. +Run the registrations and the insert in one transaction so a crash cannot leave a blog +registered with no note. **The duplicate-key error message has changed.** `DataAccess.cs` compares against the literal string @@ -452,7 +477,7 @@ on `TimeStamp` either before or after: ```sql -- now UPDATE Notes SET replyText = @replyText, DateModified = @dateModified - WHERE NoteBlogId = (SELECT BlogId FROM BlogNames WHERE BlogName = @noteBlogName) + WHERE NoteBlogId = (SELECT BlogId FROM Blogs WHERE BlogName = @noteBlogName) AND ABS(TimeStamp - @TimeStamp) <= 5 AND TypeId = (SELECT TypeId FROM NoteTypes WHERE Type = 'reply') AND (replyText IS NULL OR replyText = '' OR replyText = '.') @@ -462,18 +487,15 @@ UPDATE Notes SET replyText = @replyText, DateModified = @dateModified Rolodex's soft-delete updates need no change beyond the `WHERE` clause — they set `IsActive`, which is untouched. -### Three traps +### Two traps -**`Blogs.BlogId` is NULL on 168,202 of 188,620 rows.** Any inner join on it silently drops +**`Blogs.BlogId` is NULL on 165,887 of 198,560 rows.** Any inner join on it silently drops every blog that has never appeared in a note. Correct for engagement queries; wrong for registry listings, which need a `LEFT JOIN` or no join at all. -**12 names in `BlogNames` have no `Blogs` row.** Resolving an ID to a name through `Blogs` -will occasionally find nothing. Use `BlogNames` for names and `Blogs` for registry columns. - -**IDs are stable and must stay so.** `BlogNames.BlogId` and `NoteTypes.TypeId` are stored -in over a million `Notes` rows. Never renumber. A blog renamed upstream gets a new row, -not an edited one, unless every `Notes` reference migrates with it. +**IDs are stable and must stay so.** `Blogs.BlogId` and `NoteTypes.TypeId` are stored in +over a million `Notes` rows. Never renumber. A blog renamed upstream gets a new row, not +an edited one. The `Blogs` triggers reject both. --- @@ -481,16 +503,12 @@ not an edited one, unless every `Notes` reference migrates with it. There are no foreign keys, and the tables do not perfectly agree: -- 4 `Posts` rows name a blog with no `Blogs` row. -- 12 of the 20,430 names in `BlogNames` have no `Blogs` row. - -So a name appearing in `Notes` or `Posts` is not a guarantee that the registry knows about -it. Joins from those tables back to `Blogs` should tolerate a miss. - -The integer schema does not fix this and was not meant to. `BlogNames` is deliberately -built from `Notes` rather than from `Blogs`, precisely so that the 12 unregistered -engagers keep their IDs and their rows. Had it been built from the registry, those notes -would have been dropped by the migration's inner joins. +- 4 `Posts` rows name a blog with no `Blogs` row, so joins from `Posts` back to `Blogs` + should tolerate a miss. +- `Notes` is covered: every `RootBlogId` and `NoteBlogId` resolves to a `Blogs` row. + `retire-blognames.sql` checked this before committing, and `RegisterBlog` keeps it true. + Before 2026-09-28, 12 to 17 note participants had no registry row. They now have stub + rows. --- @@ -542,11 +560,10 @@ Crawler bookkeeping. Rolodex ignores all of these. `DataAccess.cs` joins on it to decide what to collect: ```sql --- shape only; the ported GetBlogs joins Blogs directly on BlogId and needs no BlogNames hop -SELECT bn.BlogName, count(*) +-- shape only +SELECT b.BlogName, count(*) FROM Notes n - JOIN Blogs b ON b.BlogId = n.NoteBlogId - JOIN BlogNames bn ON bn.BlogId = n.NoteBlogId + JOIN Blogs b ON b.BlogId = n.NoteBlogId WHERE b.IsActive = @isActive AND ... ``` @@ -644,7 +661,7 @@ handled: SELECT 'Blogs', COUNT(*) FROM Blogs UNION ALL SELECT 'Posts', COUNT(*) FROM Posts UNION ALL SELECT 'Notes', COUNT(*) FROM Notes -UNION ALL SELECT 'BlogNames', COUNT(*) FROM BlogNames; +UNION ALL SELECT 'Blogs with a BlogId', COUNT(*) FROM Blogs WHERE BlogId IS NOT NULL; -- note type mix (joins NoteTypes; Notes.Type no longer exists) SELECT t.Type, COUNT(*) @@ -668,8 +685,9 @@ SELECT COUNT(*) FROM ( SELECT COUNT(*) FROM Posts p WHERE NOT EXISTS (SELECT 1 FROM Blogs b WHERE b.BlogName = p.BlogName); -SELECT COUNT(*) FROM BlogNames bn -WHERE NOT EXISTS (SELECT 1 FROM Blogs b WHERE b.BlogName = bn.BlogName); +-- note participants with no Blogs.BlogId (expect 0; anything else is the pre-2026-09-28 drift) +SELECT COUNT(*) FROM (SELECT DISTINCT NoteBlogId AS Id FROM Notes) n +WHERE NOT EXISTS (SELECT 1 FROM Blogs b WHERE b.BlogId = n.Id); -- space by object, to see where the file actually goes SELECT name, SUM(pgsize)/1024/1024 AS mb diff --git a/normalize-notes.sql b/normalize-notes.sql index 5602564..c5db9d7 100644 --- a/normalize-notes.sql +++ b/normalize-notes.sql @@ -2,6 +2,10 @@ -- Replaces the repeated blog-name and type TEXT in Notes with integer IDs. -- Reduces TL.db from ~207 MB to ~148 MB (-29%). -- +-- SUPERSEDED IN PART, 2026-09-28: the BlogNames table this creates is no longer the +-- ID authority. Run retire-blognames.sql straight after this one; it moves the IDs +-- into Blogs.BlogId and drops BlogNames. The current app code assumes both have run. +-- -- THIS IS A BREAKING SCHEMA CHANGE. There is no compatibility layer. Every -- query in URLNotesGrabberCORE and Rolodex that names Notes.RootBlogName, -- Notes.NoteBlogName or Notes.Type stops working the moment this runs, and diff --git a/retire-blognames.sql b/retire-blognames.sql new file mode 100644 index 0000000..8d0781f --- /dev/null +++ b/retire-blognames.sql @@ -0,0 +1,143 @@ +-- retire-blognames.sql +-- Makes Blogs.BlogId the only ID authority for Notes and retires the BlogNames table. +-- +-- WHY: normalize-notes.sql (2026-08-07) put the IDs in BlogNames and copied them into +-- Blogs.BlogId once. Nothing kept the copy current: by 2026-09-28, 12,238 blogs first seen +-- in a note after the migration had a BlogNames ID but Blogs.BlogId = NULL, so every query +-- joining Notes to Blogs on BlogId (GetBlogs and friends) silently skipped them -- 23,148 +-- notes. Two copies of one ID drift; this leaves one. +-- +-- What it does: +-- 1. Gives every BlogNames name a Blogs row (17 had none), carrying its ID over. +-- 2. Copies the ID onto every Blogs row that is missing it. IDs are never renumbered -- +-- they are stored in 1.18M Notes rows. +-- 3. Proves every Notes ID resolves through Blogs before anything is dropped. +-- 4. Makes ix_Blogs_BlogId UNIQUE. +-- 5. Drops BlogNames. No compatibility view: any other app that still names it gets +-- "no such table: BlogNames" and must port to Blogs.BlogId (see TL.db.md). +-- 6. Adds triggers that stop a Blogs row holding a BlogId from being deleted, renamed or +-- renumbered -- the guarantees BlogNames gave by never being touched. +-- +-- DateModified is NOT moved: assigning an ID is bookkeeping, not a content change. The 17 +-- new stub rows get DateAdded/DateModified/DateCreated = now, as AddBlog would give them. +-- +-- Runs after normalize-notes.sql. A backup from before 2026-08-07 needs both, in order. +-- +-- HOW TO RUN: +-- 1. Stop every app that uses TL.db. Pause NextCloud sync. +-- 2. Back up TL.db: sqlite3 TL.db ".backup 'TL pre-retire-blognames.db'" +-- 3. sqlite3 -bail TL.db < retire-blognames.sql +-- -bail matters: a failed check aborts before COMMIT and nothing is changed. +-- In DB Browser, Execute SQL stops at the first error; then Revert Changes. +-- 4. Run the build of URLNotesGrabberCORE that no longer uses BlogNames. An older +-- build fails every AddNote with "no such table: BlogNames". + +PRAGMA foreign_keys = off; + +BEGIN; + +-- Every check inserts one count here; the CHECK aborts the script on anything but 0. +CREATE TEMP TABLE MustBeZero (Check_ TEXT, n INTEGER CHECK (n = 0)); + +-------------------------------------------------------------------------- +-- STEP 0: the two copies must not disagree anywhere they are both set +-------------------------------------------------------------------------- +INSERT INTO MustBeZero +SELECT 'Blogs.BlogId differs from BlogNames', COUNT(*) + FROM Blogs b JOIN BlogNames bn ON bn.BlogName = b.BlogName + WHERE b.BlogId <> bn.BlogId; + +INSERT INTO MustBeZero +SELECT 'Blogs.BlogId unknown to BlogNames', COUNT(*) + FROM Blogs b + WHERE b.BlogId IS NOT NULL + AND NOT EXISTS (SELECT 1 FROM BlogNames bn WHERE bn.BlogId = b.BlogId AND bn.BlogName = b.BlogName); + +-------------------------------------------------------------------------- +-- STEP 1: a Blogs row for every name Notes points at +-------------------------------------------------------------------------- +-- No IsActive in the column list: it is not ours to write (defaults to live). +INSERT INTO Blogs (BlogName, DateAdded, DateModified, DateCreated, BlogId) +SELECT bn.BlogName, + strftime('%Y-%m-%d %H:%M:%S', 'now', 'localtime'), + strftime('%Y-%m-%d %H:%M:%S', 'now', 'localtime'), + strftime('%Y-%m-%d %H:%M:%S', 'now', 'localtime'), + bn.BlogId + FROM BlogNames bn + WHERE NOT EXISTS (SELECT 1 FROM Blogs b WHERE b.BlogName = bn.BlogName); + +-------------------------------------------------------------------------- +-- STEP 2: backfill the IDs Blogs never received +-------------------------------------------------------------------------- +UPDATE Blogs + SET BlogId = (SELECT bn.BlogId FROM BlogNames bn WHERE bn.BlogName = Blogs.BlogName) + WHERE BlogId IS NULL + AND BlogName IN (SELECT BlogName FROM BlogNames); + +-------------------------------------------------------------------------- +-- STEP 3: prove Blogs now holds exactly what BlogNames held +-------------------------------------------------------------------------- +INSERT INTO MustBeZero +SELECT 'BlogNames pair missing from Blogs', COUNT(*) + FROM BlogNames bn + WHERE NOT EXISTS (SELECT 1 FROM Blogs b WHERE b.BlogId = bn.BlogId AND b.BlogName = bn.BlogName); + +INSERT INTO MustBeZero +SELECT 'Blogs IDs vs BlogNames rows', (SELECT COUNT(*) FROM Blogs WHERE BlogId IS NOT NULL) - (SELECT COUNT(*) FROM BlogNames); + +INSERT INTO MustBeZero +SELECT 'Notes.RootBlogId unresolved', COUNT(*) + FROM (SELECT DISTINCT RootBlogId AS Id FROM Notes) n + WHERE NOT EXISTS (SELECT 1 FROM Blogs b WHERE b.BlogId = n.Id); + +INSERT INTO MustBeZero +SELECT 'Notes.NoteBlogId unresolved', COUNT(*) + FROM (SELECT DISTINCT NoteBlogId AS Id FROM Notes) n + WHERE NOT EXISTS (SELECT 1 FROM Blogs b WHERE b.BlogId = n.Id); + +-------------------------------------------------------------------------- +-- STEP 4: one row per ID +-------------------------------------------------------------------------- +-- UNIQUE still allows the NULLs on the ~168k blogs that have never appeared in a note. +DROP INDEX ix_Blogs_BlogId; +CREATE UNIQUE INDEX ix_Blogs_BlogId ON Blogs (BlogId); + +-------------------------------------------------------------------------- +-- STEP 5: BlogNames goes +-------------------------------------------------------------------------- +DROP TABLE BlogNames; + +-------------------------------------------------------------------------- +-- STEP 6: what BlogNames guaranteed by never being written +-------------------------------------------------------------------------- +-- A deleted row would orphan its notes, and MAX(BlogId) + 1 in RegisterBlog could then +-- hand the same ID to a different blog. Remove a blog with IsActive = 0 instead. +CREATE TRIGGER trg_Blogs_BlogId_NoDelete +BEFORE DELETE ON Blogs +WHEN OLD.BlogId IS NOT NULL +BEGIN + SELECT RAISE(ABORT, 'Blogs row has a BlogId that Notes points at; set IsActive = 0 instead of deleting'); +END; + +-- A blog renamed upstream is a new blog to Tumblr's API and gets a new row. Editing the name +-- in place would re-attribute every note to it; changing the ID would orphan them. +CREATE TRIGGER trg_Blogs_BlogId_Immutable +BEFORE UPDATE OF BlogId, BlogName ON Blogs +WHEN OLD.BlogId IS NOT NULL + AND (NEW.BlogId IS NOT OLD.BlogId OR NEW.BlogName IS NOT OLD.BlogName) +BEGIN + SELECT RAISE(ABORT, 'BlogId and BlogName are fixed once a blog has a BlogId; Notes rows point at it'); +END; + +DROP TABLE temp.MustBeZero; + +COMMIT; + +-------------------------------------------------------------------------- +-- VERIFY +-------------------------------------------------------------------------- +-- SELECT COUNT(*) FROM sqlite_master WHERE name = 'BlogNames'; -- expect: 0 +-- SELECT COUNT(*) FROM Blogs WHERE BlogId IS NOT NULL; -- expect: the old BlogNames row count +-- SELECT sql FROM sqlite_master WHERE name = 'ix_Blogs_BlogId'; -- expect: CREATE UNIQUE INDEX +-- SELECT name FROM sqlite_master WHERE type = 'trigger'; -- expect: both triggers +-- PRAGMA integrity_check; -- expect: ok diff --git a/verify-db-schema.sql b/verify-db-schema.sql index 43a4dea..62b533a 100644 --- a/verify-db-schema.sql +++ b/verify-db-schema.sql @@ -82,7 +82,8 @@ WITH expected(tbl, col, alter_stmt) AS ( -- Blogs.BlogId (2026-08-07) is the single-hop join key into Notes. Deliberately NOT -- auto-fixable: an added-but-empty BlogId makes every engagement join return zero -- rows silently, which is worse than the hard error a missing column gives. - ('Blogs','BlogId', 'MANUAL REVIEW - see query 1d: run normalize-notes.sql'), + -- Since 2026-09-28 it is the only blog-ID authority (query 1e). + ('Blogs','BlogId', 'MANUAL REVIEW - see queries 1d/1e: run normalize-notes.sql, then retire-blognames.sql'), -- Notes (base columns: manual review if missing) -- Integer IDs since 2026-08-07. RootBlogName/NoteBlogName/Type are GONE, not renamed @@ -101,11 +102,11 @@ WITH expected(tbl, col, alter_stmt) AS ( -- placeholder. EnsureReplyTextColumnExists in DataAccess.cs adds it the same way. ('Notes','replyText', 'ALTER TABLE Notes ADD COLUMN replyText TEXT;'), - -- BlogNames / NoteTypes (the lookup tables Notes resolves its IDs through, 2026-08-07). - -- Not auto-fixable: an empty BlogNames does not mean "add the table", it means the + -- NoteTypes (the lookup table Notes resolves TypeId through, 2026-08-07). + -- Not auto-fixable: an empty NoteTypes does not mean "add the table", it means the -- Notes rows have nothing to resolve against. Rebuild with normalize-notes.sql. - ('BlogNames','BlogId', 'MANUAL REVIEW - see query 1d: run normalize-notes.sql'), - ('BlogNames','BlogName', 'MANUAL REVIEW - see query 1d: run normalize-notes.sql'), + -- BlogNames is not listed: it was dropped on 2026-09-28. Query 1e reports a file + -- that still has it. ('NoteTypes','TypeId', 'MANUAL REVIEW - see query 1d: run normalize-notes.sql'), ('NoteTypes','Type', 'MANUAL REVIEW - see query 1d: run normalize-notes.sql'), @@ -123,7 +124,6 @@ actual(tbl, col) AS ( SELECT 'Posts', name FROM pragma_table_info('Posts') UNION ALL SELECT 'Blogs', name FROM pragma_table_info('Blogs') UNION ALL SELECT 'Notes', name FROM pragma_table_info('Notes') - UNION ALL SELECT 'BlogNames', name FROM pragma_table_info('BlogNames') UNION ALL SELECT 'NoteTypes', name FROM pragma_table_info('NoteTypes') UNION ALL SELECT 'DailyAPICount', name FROM pragma_table_info('DailyAPICount') UNION ALL SELECT 'ApiKeyPoolState', name FROM pragma_table_info('ApiKeyPoolState') @@ -146,7 +146,7 @@ ORDER BY (e.alter_stmt LIKE 'ALTER%') DESC, e.tbl, e.col; -- 1b. MISSING TABLES: expected tables that don't exist at all in this DB. -- Zero rows = good. WITH expected_tables(tbl) AS ( - VALUES ('Posts'),('Blogs'),('Notes'),('BlogNames'),('NoteTypes'),('DailyAPICount'), + VALUES ('Posts'),('Blogs'),('Notes'),('NoteTypes'),('DailyAPICount'), ('ApiKeyPoolState'),('ApiKeyPoolMeta') ) SELECT et.tbl AS missing_table @@ -180,7 +180,6 @@ WITH expected(tbl, col) AS ( ('Notes','RootBlogId'),('Notes','PostID'),('Notes','NoteBlogId'),('Notes','TimeStamp'), ('Notes','TypeId'),('Notes','DatetimeCrawled'),('Notes','DateModified'),('Notes','DateCreated'), ('Notes','replyText'),('Notes','IsActive'), - ('BlogNames','BlogId'),('BlogNames','BlogName'), ('NoteTypes','TypeId'),('NoteTypes','Type'), ('DailyAPICount','Date'),('DailyAPICount','APICount'), ('ApiKeyPoolState','KeyName'),('ApiKeyPoolState','RetryUntil'), @@ -190,7 +189,6 @@ actual(tbl, col) AS ( SELECT 'Posts', name FROM pragma_table_info('Posts') UNION ALL SELECT 'Blogs', name FROM pragma_table_info('Blogs') UNION ALL SELECT 'Notes', name FROM pragma_table_info('Notes') - UNION ALL SELECT 'BlogNames', name FROM pragma_table_info('BlogNames') UNION ALL SELECT 'NoteTypes', name FROM pragma_table_info('NoteTypes') UNION ALL SELECT 'DailyAPICount', name FROM pragma_table_info('DailyAPICount') UNION ALL SELECT 'ApiKeyPoolState', name FROM pragma_table_info('ApiKeyPoolState') @@ -209,7 +207,7 @@ ORDER BY a.tbl, a.col; -- -- This is the one failure SECTION 2 cannot fix. Notes.RootBlogName / -- NoteBlogName / Type were replaced by RootBlogId / NoteBlogId / TypeId --- resolving through BlogNames and NoteTypes -- a data migration, not an +-- resolving through (then) BlogNames and NoteTypes -- a data migration, not an -- ADD COLUMN. There is no compatibility view, so the current code fails -- outright ("no such column: RootBlogId") against such a file. -- @@ -223,6 +221,28 @@ WHERE lower(name) IN ('rootblogname','noteblogname','type') HAVING COUNT(*) > 0; +-- 1e. BLOGNAMES NOT RETIRED: a backup from between 2026-08-07 and 2026-09-28, when +-- BlogNames still held the IDs and Blogs.BlogId was an +-- unmaintained copy. Zero rows = good. +-- +-- The current code resolves every Notes ID through Blogs.BlogId and never writes +-- BlogNames, so against such a file new blogs get IDs that can collide with +-- BlogNames' and every blog missing from Blogs.BlogId stays invisible to GetBlogs. +-- +-- Fix: back up, then run retire-blognames.sql (after normalize-notes.sql if 1d +-- also reported). It checks itself and changes nothing if a check fails. +SELECT 'BlogNames still exists (' || type || ') -- run retire-blognames.sql' AS blognames_not_retired +FROM sqlite_master +WHERE lower(name) = 'blognames' +UNION ALL +SELECT 'Blogs.BlogId is not UNIQUE -- run retire-blognames.sql' +WHERE NOT EXISTS (SELECT 1 FROM pragma_index_list('Blogs') WHERE name = 'ix_Blogs_BlogId' AND "unique" = 1) +UNION ALL +SELECT 'BlogId guard trigger missing: ' || t.name || ' -- run retire-blognames.sql' +FROM (SELECT 'trg_Blogs_BlogId_NoDelete' AS name UNION ALL SELECT 'trg_Blogs_BlogId_Immutable') t +WHERE NOT EXISTS (SELECT 1 FROM sqlite_master m WHERE m.type = 'trigger' AND m.name = t.name); + + -- ============================================================================ -- SECTION 2 -- FIX (opt-in, additive only) -- @@ -233,8 +253,8 @@ HAVING COUNT(*) > 0; -- subset. These are the 8 additive migration columns and nothing else; the -- likes high-water-mark reset is intentionally NOT included. -- --- Nothing here addresses query 1d. The Notes integer schema is a data migration --- (normalize-notes.sql) and cannot be reached by adding columns. +-- Nothing here addresses queries 1d or 1e. Those are data migrations +-- (normalize-notes.sql, retire-blognames.sql) and cannot be reached by adding columns. -- ============================================================================ -- ALTER TABLE Posts ADD COLUMN PostType TEXT;