Files
URLNotesGrabberCore/URLNotesGrabberCORE/TL.db.md
T
jimandClaude Opus 5.5 8fe2ffeb96 fix(db): make Blogs.BlogId the only blog ID and drop BlogNames
Blogs.BlogId was a one-time copy of BlogNames and nothing kept it
current: 12,238 blogs first seen after 2026-08-07 had a BlogNames ID
but a NULL Blogs.BlogId, so GetBlogs' join on BlogId silently skipped
them and their 23,148 notes.

- retire-blognames.sql: stub Blogs rows for the 17 unregistered note
  participants, backfill IDs (none renumbered), make ix_Blogs_BlogId
  UNIQUE, drop BlogNames, and add triggers that stop a Blogs row with
  a BlogId from being deleted, renamed or renumbered
- AddNote registers both blogs via RegisterBlog (Blogs row + MAX+1 ID)
  and every query resolves names through Blogs instead of BlogNames
- verify-db-schema.sql reports a DB that still has BlogNames (1e)
- Update TL.db.md, AGENTS.md and the DB Browser saved queries

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-28 11:33:04 -05:00

31 KiB

TL.db — schema notes

The SQLite database behind URLNotesGrabberCORE and its sibling crawlers, and the one Rolodex reads.

Everything below was read out of the live file, not inferred from code. Counts are as of 2026-08-07; re-run the queries at the bottom to refresh them.

  • Journal mode: WAL — TL.db-wal and TL.db-shm live beside the file and are part of the database. Copying TL.db alone gives you whatever was last checkpointed, not the current state.
  • Page size: 4096. File size: 148 MB.

⚠ Breaking change, 2026-08-07: Notes holds integer IDs, not names

Notes.RootBlogName, Notes.NoteBlogName and Notes.Type no longer exist. They are now RootBlogId, NoteBlogId and TypeId, resolved through the new BlogNames and NoteTypes tables. Any query naming the old columns fails outright.

There is no compatibility view. See porting to the integer schema for the old-to-new translation of every query shape the applications use.

Applied by ../normalize-notes.sql, which took the file from 207 MB to 148 MB. An earlier change the same day (../shrink-db.sql) took it from 267 MB to 207 MB.

⚠ Breaking change, 2026-09-28: BlogNames is gone; Blogs.BlogId is the only ID authority

The IDs in Notes used to live in a BlogNames table, with a copy in Blogs.BlogId. Nothing kept the copy current, so by 2026-09-28 12,238 blogs first seen after the migration had Blogs.BlogId = NULL. Every Notes-to-Blogs join on BlogId silently skipped them and their 23,148 notes, which kept them out of GetBlogs.

../retire-blognames.sql fixed this by giving every note participant a Blogs row, backfilling the IDs (none renumbered), making ix_Blogs_BlogId unique, and dropping BlogNames. There is no compatibility view: any query naming it fails with no such table: BlogNames. Two triggers now protect the IDs.

Porting an app: replace BlogNames with Blogs everywhere. The columns you used, BlogId and BlogName, exist there with the same meaning. A name lookup (SELECT BlogId FROM Blogs WHERE BlogName = ?) is a primary-key probe, and an ID lookup or join (JOIN Blogs b ON b.BlogId = n.NoteBlogId) uses the unique ix_Blogs_BlogId. Every ID in Notes resolves to exactly one Blogs row. Blogs.BlogId is no longer a stale copy, so any code or docs that distrust it can drop that caveat. Never write BlogId or BlogName on a row that has an ID, and never delete such a row: the triggers reject all three. See Blogs.


The three content tables

Table Rows What it is
Blogs 198,560 The crawl registry, one row per known blog plus crawl-state flags. Also the ID authority for blogs in Notes
Posts 22,468 Stored post content. Only 3,867 blogs actually have any
Notes 1,234,830 The engagement graph: NoteBlogId acted on (RootBlogId, PostID)

(Blogs and Notes counts as of 2026-09-28; the rest as of 2026-08-07.)

…supported by one lookup table that exists only to keep Notes small:

Table Rows What it is
NoteTypes 5 TypeId ⇄ Type. like, reblog, reply, posted, post_attribution

The engagement graph is the interesting part. 20,311 distinct blogs appear as engagers — far more than the 3,867 that have stored posts — which is what makes this a social graph rather than a post archive. Only 2,771 blogs appear as the root of a note.

Blogs

CREATE TABLE "Blogs" (
    "BlogName"              TEXT,
    "HasBeenOutput"         INTEGER DEFAULT 0,
    "IsActive"              INTEGER DEFAULT 1,
    "DateAdded"             TEXT NOT NULL DEFAULT '12/24/25',
    "ByLikes"               INTEGER NOT NULL DEFAULT 0,
    "LikesPulled"           INTEGER NOT NULL DEFAULT 0,
    "LikesCursor"           INTEGER DEFAULT 0,
    "DateModified"          TEXT,
    "DateCreated"           TEXT,
    LikesNewestTimestamp    INTEGER DEFAULT 0,
    LikesLastRefreshed      INTEGER DEFAULT 0,
    LikesLastNewCount       INTEGER DEFAULT 0,
    TTFolderPath            TEXT,
    BlogId                  INTEGER,
    PRIMARY KEY("BlogName")
);

CREATE UNIQUE INDEX ix_Blogs_BlogId ON Blogs (BlogId);

CREATE TRIGGER trg_Blogs_BlogId_NoDelete        -- no DELETE of a row that has a BlogId
CREATE TRIGGER trg_Blogs_BlogId_Immutable       -- no change to its BlogId or BlogName

BlogName is the primary key, so it is the only indexed way in by name. There is no index on any flag or date. Filtering or sorting on those scans the whole table, which is affordable here and is not on Notes.

BlogId is the ID that Notes.RootBlogId and Notes.NoteBlogId store, and Blogs is the only place it lives (since 2026-09-28; see the banner at the top). The join to Notes is one integer hop on the unique index:

FROM Blogs B JOIN Notes N ON N.NoteBlogId = B.BlogId

Every blog that appears in Notes has a Blogs row with a BlogId. AddNote guarantees it through RegisterBlog, which runs in the note's own transaction:

INSERT OR IGNORE INTO Blogs (BlogName, DateAdded, DateModified, DateCreated)
VALUES (@name, @now, @now, @now);
UPDATE Blogs SET BlogId = (SELECT IFNULL(MAX(BlogId), 0) + 1 FROM Blogs)
 WHERE BlogName = @name AND BlogId IS NULL;
  • Unlike AddBlog, this does not skip names containing deact. A note by a deactivated blog still needs an ID, so such blogs now get registry rows too, with the usual defaults (HasBeenOutput = 0, IsActive left at its default).
  • Assigning a BlogId is bookkeeping, so it does not move DateModified.
  • MAX(BlogId) + 1 is safe only because an ID can never be freed. The two triggers see to that: deleting a row that has a BlogId, or changing its BlogId or BlogName, aborts. Remove a blog with IsActive = 0 instead. A blog renamed upstream gets a new row. Rows with no BlogId can still be deleted or renamed freely.
  • INSERT OR REPLACE on Blogs gets around the delete trigger (SQLite does not fire delete triggers for REPLACE unless recursive_triggers is on), and it would wipe the BlogId. It was already forbidden because it resets IsActive. Do not use it.

BlogId is NULL on 165,887 of 198,560 rows, every blog that has never appeared in a note. That is the large majority, and it is not an error: the registry is far bigger than the engagement graph. An inner join on BlogId therefore silently drops those blogs, which is usually what you want for engagement queries and is wrong for registry listings.

Flag distribution: IsActive = 1 on 188,601 of 188,620 rows, HasBeenOutput = 1 on 5,059, ByLikes = 1 on 2. IsActive carries a second meaning as of Rolodex — see Blogs.IsActive below.

The columns after DateCreated were added later by ALTER TABLE, which is why they carry no quoting in the stored DDL. That is the normal way this schema grows, and BlogId is the newest example.

DateAdded is not written consistently. 170,677 rows hold ISO yyyy-MM-dd HH:mm:ss; 17,943 hold US-format M/d/yy from a bulk import. As text those two sort into different parts of the table, so anything ordering or range-filtering on this column has to normalise first — see DateSql in Rolodex.

Posts

CREATE TABLE "Posts" (
    "BlogName"              TEXT,
    "PostID"                INTEGER,
    "HasNotesGathered"      INTEGER DEFAULT 0,
    "reblogURL"             TEXT,
    "NotFound"              INTEGER DEFAULT 0,
    "PostDate"              TEXT,
    "NotesGatheredDateTime" INTEGER NOT NULL DEFAULT 1729746000,
    "HasImage"              INTEGER NOT NULL DEFAULT 0,
    "PostURL"               TEXT,
    "Slug"                  TEXT,
    "ReblogKey"             TEXT,
    "ReblogName"            TEXT,
    "Summary"               TEXT,
    "Quote"                 TEXT,
    "Body"                  TEXT,
    "Tags"                  TEXT,
    "Link"                  TEXT,
    "PhotoURL"              TEXT,
    "PhotoCaption"          TEXT,
    "DownloadedFiles"       TEXT,
    "AudioCaption"          TEXT,
    "Question"              TEXT,
    "Answer"                TEXT,
    "Title"                 TEXT,
    "ByLikes"               INTEGER NOT NULL DEFAULT 0,
    "RootBlogName"          TEXT,
    "RootURL"               TEXT,
    "DateModified"          TEXT,
    "DateCreated"           TEXT,
    PostType                TEXT,
    PRIMARY KEY("BlogName","PostID")
);

The key is (BlogName, PostID), not PostID. This matters more than it looks: 345 post IDs exist under more than one blog, so an ID on its own is both ambiguous and unindexed. Any lookup should carry the blog name, and a batch lookup should group by blog so it stays on the leading column of the key.

Notable:

  • PostType is now mostly populated: 20,679 of 22,468 rows, leaving 1,789 NULL. This reverses what earlier revisions of this document said — the column really was empty on every row, and something has since started writing it. Anything that treated it as permanently unset, or derived the type from post content instead, should be re-examined against the live data. Rolodex still derives it.
  • PostDate is yyyy-MM-dd HH:mm:ss GMT — UTC, as the Tumblr API sends it, unlike the local-time DateCreated/DateModified. Text-file exports may carry RFC 1123 (Fri, 14 Feb 2025 15:20:09 GMT); every write path runs PostDates.Normalize to convert it, and ../normalize-postdate.sql fixed the 4 rows written before that.
  • HasImage = 1 on 14,026 rows. It records that the post had a picture, not that a usable URL was kept, so it is not a reliable predictor that anything will render.
  • PhotoURL is largely unused; in practice the image markup lives inside Body.
  • NotFound = 1 on 4,712 rows — posts that have since been deleted upstream.
  • The content columns (Body, Quote, Question, Answer, …) are the heavy ones. List views should not select them.

Notes

CREATE TABLE Notes (
    RootBlogId      INTEGER NOT NULL,
    PostID          INTEGER NOT NULL,
    NoteBlogId      INTEGER NOT NULL,
    TimeStamp       INTEGER NOT NULL,
    TypeId          INTEGER NOT NULL,
    replyText       TEXT,
    DatetimeCrawled TEXT,
    DateModified    TEXT,
    DateCreated     TEXT,
    IsActive        INTEGER NOT NULL DEFAULT 1,
    PRIMARY KEY (RootBlogId, PostID, TimeStamp, TypeId, NoteBlogId)
) WITHOUT ROWID;

CREATE INDEX ix_Notes_NoteBlogId ON Notes (NoteBlogId);

Integer IDs since 2026-08-07 — this is the breaking change. RootBlogName, NoteBlogName and Type are gone, replaced by RootBlogId, NoteBlogId and TypeId. Resolve blog IDs through Blogs.BlogId and types through NoteTypes. The old names were text repeated on 1.18 million rows, in the table and in every index over it; the swap took the file from 207 MB to 148 MB.

The primary key column order is deliberately unchanged, so the leading-prefix access patterns callers already depend on still hold: (RootBlogId) and (RootBlogId, PostID) remain cheap prefixes, exactly as (RootBlogName) and (RootBlogName, PostID) were.

Two nulls-and-defaults differences from the old DDL, both intentional:

  • replyText and DatetimeCrawled no longer carry column defaults. The old table defaulted them to '.' and '2/12/26 12am', which is how 1.1M rows acquired placeholder values nobody wrote. New rows now get NULL unless a writer supplies something. The crawler names both columns explicitly, so its behaviour is unchanged.
  • The five key columns are now NOT NULL. They always were in practice.

WITHOUT ROWID, since earlier the same day. The rows live in the primary key's b-tree rather than in a rowid table with a separate key index beside it. Two consequences matter before adding an index here:

  • There is no rowid on this table. SELECT rowid FROM Notes is an error, and no code in any of the three apps relied on it.
  • A secondary index carries the whole five-column primary key as its row reference instead of a compact rowid, so indexes here are expensive — though far less so than before, now that the key is five integers rather than three integers and two strings. ix_Notes_NoteBlogId costs 27 MB; its text predecessor cost 58 MB.

Notes_idx_06e01ae3 on TimeStamp DESC was dropped at the same time. It cost 14 MB as a rowid index and would have cost 58 MB after the conversion. It was worth neither: the crawler's only TimeStamp filter (>= 1535778000) excludes 786 rows of 1.18M, Rolodex's default Notes sort carries a three-column tiebreaker that forces a full sort regardless, and the reply-matching UPDATE uses ABS(TimeStamp - ?) <= 5, which no index on TimeStamp can serve. The one path that got slower is Rolodex's Notes page with a date-range filter: 60 ms to 164 ms.

See ../shrink-db.sql for the full rationale and the applied result.

DatetimeCrawled is NULL on 1,148,077 rows, and that is the honest value. Those rows previously stored the literal string '2/12/26 12am' — this column's own DDL default, written as a bulk backfill placeholder rather than as a crawl time. They were set to NULL on 2026-08-07, which is what consumers already displayed them as: the string parses as a date in neither format this schema writes.

Note the trap: the DEFAULT '2/12/26 12am' clause is still in the DDL above. Any INSERT that omits this column writes the placeholder straight back. The crawler names it explicitly on every insert, so nothing reintroduces it today, but a new writer that forgets to would — which is why consumers should keep treating an unparseable value here as "unknown" rather than assuming NULL is now the only such marker.

One row per engagement event. TimeStamp is unix seconds — unlike every date column elsewhere in the schema, which are text.

At 1.18M rows this is the table that dictates how the whole database has to be queried:

  • Nothing should run an unbounded SELECT or a bare COUNT(*) here. A count scans the lot on every call.
  • The only fast access paths are the primary key's leading columns (RootBlogId, then PostID) and ix_Notes_NoteBlogId on NoteBlogId. "Notes received by a blog" and "notes given by a blog" are both cheap; almost nothing else is.
  • Every ordering here is a full sort of whatever the filters leave, TimeStamp included. Filter first, then sort.
  • replyText is '.' on 1,167,464 rows — only reply notes carry real text. Those dots are inherited from the old column default; new rows get NULL instead.

Resolve IDs by filtering Blogs, not by scanning Notes. A name predicate on Blogs is a primary-key probe, so pushing it there costs nothing and lets the Notes index do the work:

-- good: Blogs resolves the name, then the index is searched
SELECT * FROM Notes
 WHERE NoteBlogId = (SELECT BlogId FROM Blogs WHERE BlogName = ?);

-- also good, same plan
SELECT n.* FROM Notes n
  JOIN Blogs b ON b.BlogId = n.NoteBlogId
 WHERE b.BlogName = ?;

NoteTypes

CREATE TABLE NoteTypes (
    TypeId INTEGER PRIMARY KEY,
    Type   TEXT NOT NULL UNIQUE
);
TypeId Type Rows Share
1 like 945,167 79.9%
2 reblog 219,203 18.5%
3 reply 15,345 1.3%
4 posted 2,617 0.2%
5 post_attribution 1 —

The set is fixed in practice, but it is a table rather than a CHECK constraint so that adding a type is an INSERT and not a schema migration. The IDs above are stored in Notes and must not be reassigned.

Five rows means the lookup is effectively free; write t.Type = 'reblog' and let SQLite resolve it, or hardcode the ID if you prefer — both are fine, but hardcoding ties your code to this table's contents, so prefer the join in anything long-lived.


Porting to the integer schema

Written for the 2026-08-07 change, and updated for 2026-09-28: wherever this section once said BlogNames, it now says Blogs. BlogNames no longer exists.

Everything here was checked against the live 148 MB file. There were 14 affected call sites in DataAccess.cs and 16 in RolodexRepository.cs. TumblThree needs no changes — its single statement touches Blogs.IsActive and BlogName only.

DataAccess.cs is ported. All 14 sites now read the integer schema, AddNote registers names and types before inserting, and verify-db-schema.sql reports a pre-migration file rather than letting the app fail on it. RolodexRepository.cs lives in the Rolodex repository and is not covered by that work. One site was dropped rather than translated: the LEFT JOIN Notes in GetPosts selected nothing and was collapsed by the query's own GROUP BY, so it could not affect the result.

Column mapping

Was Is now Resolve via
Notes.RootBlogName Notes.RootBlogId Blogs.BlogId → .BlogName
Notes.NoteBlogName Notes.NoteBlogId Blogs.BlogId → .BlogName
Notes.Type Notes.TypeId NoteTypes.TypeId → .Type
ix_NoteBlogName01 ix_Notes_NoteBlogId —

PostID, TimeStamp, replyText, DatetimeCrawled, DateModified, DateCreated and IsActive are unchanged.

Filtering by a blog name

-- was
WHERE NoteBlogName = @Name

-- now, either form; both search ix_Notes_NoteBlogId after a primary-key lookup
WHERE NoteBlogId = (SELECT BlogId FROM Blogs WHERE BlogName = @Name)
-- or
JOIN Blogs b ON b.BlogId = n.NoteBlogId WHERE b.BlogName = @Name

Measured at 73 ms against 63 ms for the old text form on the busiest blog, when the lookup was still BlogNames. The extra hop is one index probe and does not show.

Joining Notes to Blogs

This is the join to get right; it is the most common shape in both applications.

-- was
FROM Blogs B INNER JOIN Notes N ON N.NoteBlogName = B.BlogName

-- now: one integer hop, using Blogs.BlogId
FROM Blogs B INNER JOIN Notes N ON N.NoteBlogId = B.BlogId

Joining Notes to Posts also goes through Blogs, since Posts has only a name:

FROM Posts P
JOIN Blogs RB ON RB.BlogName = P.BlogName
JOIN Notes N  ON N.RootBlogId = RB.BlogId AND N.PostID = P.PostID

Selecting a name back out

-- was
SELECT NoteBlogName AS blogName, COUNT(*) FROM Notes ... GROUP BY NoteBlogName

-- now
SELECT b.BlogName AS blogName, COUNT(*)
  FROM Notes n JOIN Blogs b ON b.BlogId = n.NoteBlogId
 ... GROUP BY b.BlogName

Group by n.NoteBlogId instead of b.BlogName when you only need the name for display — grouping on the integer is cheaper and the name comes along for free.

Filtering by type

-- was
WHERE type IN ('reblog', 'reply', 'posted')

-- now
WHERE TypeId IN (SELECT TypeId FROM NoteTypes WHERE Type IN ('reblog','reply','posted'))
-- or, equivalently
JOIN NoteTypes t ON t.TypeId = n.TypeId WHERE t.Type IN ('reblog','reply','posted')

WHERE TypeId IN (2,3,4) also works and is marginally faster, but hardcodes this table's contents into application code. Prefer the lookup outside of hot paths.

Note the negated form needs care: type NOT IN ('reblog','reply','posted') becomes TypeId NOT IN (SELECT TypeId FROM NoteTypes WHERE Type IN (...)), which is correct only because TypeId is NOT NULL.

Inserting a note

The crawler must ensure both blogs have IDs first: run the RegisterBlog pair shown under Blogs for each name. No read-back, no round trip, and safe to run every time. Then:

INSERT OR IGNORE INTO Notes
       (RootBlogId, PostID, NoteBlogId, TimeStamp, TypeId,
        DatetimeCrawled, DateModified, DateCreated)
SELECT (SELECT BlogId FROM Blogs WHERE BlogName = @rootBlogName),
       @PostID,
       (SELECT BlogId FROM Blogs WHERE BlogName = @noteBlogName),
       @TimeStamp,
       (SELECT TypeId FROM NoteTypes WHERE Type = @Type),
       @DatetimeCrawled, @DateModified, @DateCreated;

Verified: a genuinely new note inserts, and re-running the identical statement inserts 0. Run the registrations and the insert in one transaction so a crash cannot leave a blog registered with no note.

The duplicate-key error message has changed. DataAccess.cs compares against the literal string

UNIQUE constraint failed: Notes.RootBlogName, Notes.PostID, Notes.TimeStamp, Notes.Type, Notes.NoteBlogName

at two call sites to decide whether to swallow an exception. SQLite now emits the new column names, so those comparisons no longer match and real errors will surface where they used to be silently ignored — or vice versa.

Both sites now go through IsNotesDuplicateKey in DataAccess.cs, which matches on UNIQUE constraint failed plus Notes. rather than on the column list. A literal comparison is what broke here; the next rename should not break it again.

Updating notes

Predicates translate the same way. The reply-matching update, which cannot use an index on TimeStamp either before or after:

-- now
UPDATE Notes SET replyText = @replyText, DateModified = @dateModified
 WHERE NoteBlogId = (SELECT BlogId FROM Blogs WHERE BlogName = @noteBlogName)
   AND ABS(TimeStamp - @TimeStamp) <= 5
   AND TypeId = (SELECT TypeId FROM NoteTypes WHERE Type = 'reply')
   AND (replyText IS NULL OR replyText = '' OR replyText = '.')
   AND (replyText IS NULL OR replyText <> @replyText);

Rolodex's soft-delete updates need no change beyond the WHERE clause — they set IsActive, which is untouched.

Two traps

Blogs.BlogId is NULL on 165,887 of 198,560 rows. Any inner join on it silently drops every blog that has never appeared in a note. Correct for engagement queries; wrong for registry listings, which need a LEFT JOIN or no join at all.

IDs are stable and must stay so. Blogs.BlogId and NoteTypes.TypeId are stored in over a million Notes rows. Never renumber. A blog renamed upstream gets a new row, not an edited one. The Blogs triggers reject both.


Referential integrity

There are no foreign keys, and the tables do not perfectly agree:

  • 4 Posts rows name a blog with no Blogs row, so joins from Posts back to Blogs should tolerate a miss.
  • Notes is covered: every RootBlogId and NoteBlogId resolves to a Blogs row. retire-blognames.sql checked this before committing, and RegisterBlog keeps it true. Before 2026-09-28, 12 to 17 note participants had no registry row. They now have stub rows.

The '.' placeholder convention

The crawler writes a single dot into text columns it has no value for, rather than NULL. This is the single most surprising thing about the schema and it affects every consumer.

Column '.' rows
Notes.replyText 1,167,464
Posts.Title 12,562
Posts.Body 172

Notes.replyText and Notes.DatetimeCrawled no longer carry column defaults as of the integer migration, so new note rows get NULL rather than a placeholder. The dots already in replyText were not rewritten — cleaning is still required on read.

Any query whose output reaches a human should collapse it:

NULLIF(NULLIF(SomeColumn, '.'), '') AS SomeColumn

Empty string turns up too, hence the double NULLIF. Not every column is affected — Blogs.TTFolderPath and Blogs.DateModified currently have zero dot rows — but new columns tend to acquire them, so treat cleaning as the default for any text column rendered to a user.


Supporting tables

Crawler bookkeeping. Rolodex ignores all of these.

Table Rows What it is
DailyAPICount 133 (Date TEXT PK, APICount INTEGER) — per-day API call tally against the rate limit
ApiKeyPoolState 2 (KeyName TEXT PK, RetryUntil INTEGER) — per-key backoff; RetryUntil is unix seconds
ApiKeyPoolMeta 1 (Id PK CHECK (Id = 1), LastIndex) — round-robin cursor. Singleton by check constraint
CollectRunState 1 (Id PK CHECK (Id = 1), RunCutoff, RunComplete, RunStarted, RunCompletedAt) — resume state for an interrupted collection run. Also a singleton

Blogs.IsActive — now written by two applications

IsActive has always been the crawler's work-selection flag. GetBlogs in DataAccess.cs joins on it to decide what to collect:

-- shape only
SELECT b.BlogName, count(*)
  FROM Notes n
  JOIN Blogs b ON b.BlogId = n.NoteBlogId
 WHERE b.IsActive = @isActive AND ...

Nothing inside the crawler writes it — it is an input, set from outside.

Rolodex is now one of the things that sets it. Removing a blog through the Rolodex UI runs exactly this:

UPDATE Blogs SET IsActive = 0 WHERE BlogName = ?;

Rolodex adds no column and changes no schema. It reuses this flag because the two meanings were judged to be one decision: a blog you do not want in the browsing UI is a blog you do not want to keep crawling. Removal therefore stops collection, and the Rolodex confirmation screen says so before anyone commits.

  • 1 (or absent/NULL) — live. Crawled, and visible in Rolodex.
  • 0 — removed. Not crawled, hidden from the Rolodex registry, dashboard counts and engagement rollups.

Restoring is the same UPDATE with a 1. Nothing is destroyed either way: the blog's Posts and Notes rows are never touched, and Rolodex deliberately keeps showing them under its Posts and Notes pages. Removing a blog hides the blog, not what it collected.

What other tools need to know

  1. Setting IsActive = 0 now also hides the blog from Rolodex, and setting it back to 1 makes it reappear. If another tool deactivates blogs in bulk, it is also removing them from the browsing UI — which may be exactly right, but it is no longer a crawler-only decision.
  2. Re-crawling a removed blog will not bring it back, since nothing in the crawler writes the flag. An INSERT OR REPLACE on the Blogs row would, by resetting it to the column default of 1. Prefer an UPDATE of the specific columns, or INSERT … ON CONFLICT DO UPDATE SET naming only the columns being refreshed.
  3. NULL is treated as live. The column is INTEGER DEFAULT 1 with no NOT NULL, so Rolodex reads it through COALESCE(IsActive, 1). A NULL therefore leaves the blog visible rather than stranding it outside both the registry and the removed list, where no screen could reach it. Write 0 or 1, not NULL.
  4. Backing the feature out is a configuration change, not a migration. Because there is no Rolodex-owned column, setting Rolodex__EnableBlogDeletion=false is the whole of it; there is nothing to drop. Any blogs already at IsActive = 0 simply go back to being ordinary inactive blogs.

Posts.IsActive and Notes.IsActive — present, and written from outside

The same flag extends to the two content tables, with the same meaning: 0 is removed, anything else — including NULL — is live. Both columns now exist in the live TL.db and are included in the DDL quoted above. As of 2026-08-07, Posts.IsActive = 0 on 5,900 rows and Notes.IsActive = 0 on none. Like Blogs.IsActive, they are written from outside this crawler.

On Notes the column is INTEGER NOT NULL DEFAULT 1, so a NULL cannot occur there; Posts and Blogs are laxer, which is why the predicate below still uses COALESCE.

The crawler therefore treats both as optional, and as nothing it owns:

  • It never writes them. No INSERT column list names IsActive, no UPDATE sets it, and MapPrefixToColumn — the only place a column name is chosen at runtime — cannot map to it. Re-crawling a removed post or note refreshes its content and leaves the flag at 0. There is no INSERT OR REPLACE on Posts or Notes for a default to be reset by.
  • It filters on them only when they exist. HasIsActiveColumn in DataAccess.cs asks PRAGMA table_info once per table per database path and caches the answer; the filter is COALESCE(IsActive, 1) = 1, and it is dropped entirely when the column is absent. Naming a missing column is a hard SQLite error, so this is what lets one build run against databases on both sides of the change. The cache lives for the process — adding the columns to a live database takes effect on the next run.

Every read that selects posts or notes carries the filter: GetPosts, GetReplies, GetRepliesWithMissingText, GetRepliesWithFilledText, GetAllPostTextColumns, GetAllPostsForBlog, GetPost, GetPostByIdAnyBlog, and the engagement queries that count or join Notes (GetBlogs, GetBlogsAll, GetBlogsForLikes). The one deliberate omission is the LEFT JOIN Notes in GetPosts: nothing is selected from it and it can neither add nor remove a row, so filtering it would buy nothing.

LegacyPostsDbImporter is unfiltered too — it reads a foreign legacy database whose Posts table is not this schema.

Two consequences worth stating plainly, both inherited from how Blogs.IsActive is handled:

  1. Removal hides a row; it does not freeze it. The write paths are keyed on a post the caller already selected, so an ingest or a correction run still overwrites the content of a removed post. Only selection is filtered.
  2. NULL is live. Write 0 or 1, not NULL, but a NULL leaves the row visible rather than stranding it.

Reproducing the numbers

SELECT 'Blogs', COUNT(*) FROM Blogs
UNION ALL SELECT 'Posts',     COUNT(*) FROM Posts
UNION ALL SELECT 'Notes',     COUNT(*) FROM Notes
UNION ALL SELECT 'Blogs with a BlogId', COUNT(*) FROM Blogs WHERE BlogId IS NOT NULL;

-- note type mix (joins NoteTypes; Notes.Type no longer exists)
SELECT t.Type, COUNT(*)
  FROM Notes n JOIN NoteTypes t ON t.TypeId = n.TypeId
 GROUP BY t.Type ORDER BY 2 DESC;

-- how much of the registry participates in the engagement graph
SELECT COUNT(*) FILTER (WHERE BlogId IS NOT NULL) AS with_notes,
       COUNT(*) FILTER (WHERE BlogId IS NULL)     AS without_notes
  FROM Blogs;

-- the two date shapes in Blogs.DateAdded
SELECT CASE WHEN DateAdded LIKE '____-__-__%' THEN 'ISO' ELSE 'US' END, COUNT(*)
FROM Blogs GROUP BY 1;

-- post IDs that are ambiguous without a blog name
SELECT COUNT(*) FROM (
    SELECT PostID FROM Posts GROUP BY PostID HAVING COUNT(DISTINCT BlogName) > 1);

-- rows that reference a blog the registry does not have
SELECT COUNT(*) FROM Posts p
WHERE NOT EXISTS (SELECT 1 FROM Blogs b WHERE b.BlogName = p.BlogName);

-- note participants with no Blogs.BlogId (expect 0; anything else is the pre-2026-09-28 drift)
SELECT COUNT(*) FROM (SELECT DISTINCT NoteBlogId AS Id FROM Notes) n
WHERE NOT EXISTS (SELECT 1 FROM Blogs b WHERE b.BlogId = n.Id);

-- space by object, to see where the file actually goes
SELECT name, SUM(pgsize)/1024/1024 AS mb
  FROM dbstat GROUP BY name ORDER BY SUM(pgsize) DESC;

Open the file read-only so an inspection can never disturb a running crawl:

sqlite3 "file:TL.db?mode=ro" ".schema"