Files
URLNotesGrabberCore/normalize-notes.sql
T
jimandClaude Opus 5.5 8fe2ffeb96 fix(db): make Blogs.BlogId the only blog ID and drop BlogNames
Blogs.BlogId was a one-time copy of BlogNames and nothing kept it
current: 12,238 blogs first seen after 2026-08-07 had a BlogNames ID
but a NULL Blogs.BlogId, so GetBlogs' join on BlogId silently skipped
them and their 23,148 notes.

- retire-blognames.sql: stub Blogs rows for the 17 unregistered note
  participants, backfill IDs (none renumbered), make ix_Blogs_BlogId
  UNIQUE, drop BlogNames, and add triggers that stop a Blogs row with
  a BlogId from being deleted, renamed or renumbered
- AddNote registers both blogs via RegisterBlog (Blogs row + MAX+1 ID)
  and every query resolves names through Blogs instead of BlogNames
- verify-db-schema.sql reports a DB that still has BlogNames (1e)
- Update TL.db.md, AGENTS.md and the DB Browser saved queries

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-28 11:33:04 -05:00

148 lines
6.4 KiB
PL/PgSQL

-- normalize-notes.sql
-- Replaces the repeated blog-name and type TEXT in Notes with integer IDs.
-- Reduces TL.db from ~207 MB to ~148 MB (-29%).
--
-- SUPERSEDED IN PART, 2026-09-28: the BlogNames table this creates is no longer the
-- ID authority. Run retire-blognames.sql straight after this one; it moves the IDs
-- into Blogs.BlogId and drops BlogNames. The current app code assumes both have run.
--
-- THIS IS A BREAKING SCHEMA CHANGE. There is no compatibility layer. Every
-- query in URLNotesGrabberCORE and Rolodex that names Notes.RootBlogName,
-- Notes.NoteBlogName or Notes.Type stops working the moment this runs, and
-- stays broken until those queries are rewritten. This was a deliberate choice
-- over a view-plus-triggers shim, which was measured to work but cost 194 ms ->
-- 321 ms on Rolodex's unfiltered Notes page.
--
-- TumblThree is unaffected. It touches only Blogs, and the column added to
-- Blogs here is additive.
--
-- HOW TO RUN (DB Browser for SQLite):
-- 1. Stop all three apps. Pause NextCloud sync.
-- 2. Back up TL.db.
-- 3. Execute SQL, paste this file, run. Then Write Changes.
-- 4. Tools > Compact Database (VACUUM). Nothing shrinks until this finishes.
--------------------------------------------------------------------------
-- The shape this produces
--------------------------------------------------------------------------
-- BlogNames(BlogId, BlogName) the ID authority: every name appearing in
-- Notes as either participant. 20,430 rows.
-- 12 of these have no Blogs row -- the
-- registry has never been a superset of the
-- engagement graph, and still is not.
--
-- NoteTypes(TypeId, Type) 5 rows. Fixed set, but written as a table
-- rather than a CHECK so a new type is an
-- INSERT and not a schema migration.
--
-- Notes(...Id columns...) integer FKs in place of text. WITHOUT ROWID,
-- same 5-column key in the same column order.
--
-- Blogs.BlogId NEW additive column. Lets Notes join Blogs in
-- one integer hop instead of going through
-- BlogNames and comparing text at the end.
-- NULL on the 168,202 blogs with no notes.
PRAGMA foreign_keys = off;
BEGIN;
--------------------------------------------------------------------------
-- STEP 1: the ID authority
--------------------------------------------------------------------------
CREATE TABLE BlogNames (
BlogId INTEGER PRIMARY KEY,
BlogName TEXT NOT NULL UNIQUE
);
INSERT INTO BlogNames (BlogName)
SELECT RootBlogName FROM Notes
UNION
SELECT NoteBlogName FROM Notes;
--------------------------------------------------------------------------
-- STEP 2: the type lookup
--------------------------------------------------------------------------
CREATE TABLE NoteTypes (
TypeId INTEGER PRIMARY KEY,
Type TEXT NOT NULL UNIQUE
);
-- IDs are assigned explicitly and must stay stable: they are stored in Notes.
INSERT INTO NoteTypes (TypeId, Type) VALUES
(1, 'like'),
(2, 'reblog'),
(3, 'reply'),
(4, 'posted'),
(5, 'post_attribution');
--------------------------------------------------------------------------
-- STEP 3: rebuild Notes with integer keys
--------------------------------------------------------------------------
-- Column order of the primary key is unchanged from the text version, so the
-- leading-column access patterns callers already rely on still hold:
-- (RootBlogId) and (RootBlogId, PostID) remain cheap prefixes.
CREATE TABLE NotesN (
RootBlogId INTEGER NOT NULL,
PostID INTEGER NOT NULL,
NoteBlogId INTEGER NOT NULL,
TimeStamp INTEGER NOT NULL,
TypeId INTEGER NOT NULL,
replyText TEXT,
DatetimeCrawled TEXT,
DateModified TEXT,
DateCreated TEXT,
IsActive INTEGER NOT NULL DEFAULT 1,
PRIMARY KEY (RootBlogId, PostID, TimeStamp, TypeId, NoteBlogId)
) WITHOUT ROWID;
-- Inner joins are safe here: BlogNames was just built from these very columns,
-- and NoteTypes covers all 5 values present. A row that failed to match would
-- be silently dropped, which is what the row-count check at the bottom is for.
INSERT INTO NotesN
SELECT r.BlogId, n.PostID, b.BlogId, n.TimeStamp, t.TypeId,
n.replyText, n.DatetimeCrawled, n.DateModified, n.DateCreated, n.IsActive
FROM Notes n
JOIN BlogNames r ON r.BlogName = n.RootBlogName
JOIN BlogNames b ON b.BlogName = n.NoteBlogName
JOIN NoteTypes t ON t.Type = n.Type;
DROP TABLE Notes;
ALTER TABLE NotesN RENAME TO Notes;
-- Replaces ix_NoteBlogName01. Renamed because it indexes a different column now.
CREATE INDEX ix_Notes_NoteBlogId ON Notes (NoteBlogId);
--------------------------------------------------------------------------
-- STEP 4: give Blogs the matching id
--------------------------------------------------------------------------
-- Additive: no existing column changes, so TumblThree's
-- "UPDATE Blogs SET IsActive = 0 ... WHERE BlogName = ?" is untouched.
ALTER TABLE Blogs ADD COLUMN BlogId INTEGER;
UPDATE Blogs
SET BlogId = (SELECT bn.BlogId FROM BlogNames bn WHERE bn.BlogName = Blogs.BlogName);
CREATE INDEX ix_Blogs_BlogId ON Blogs (BlogId);
COMMIT;
--------------------------------------------------------------------------
-- STEP 5: Write Changes, then Tools > Compact Database
--------------------------------------------------------------------------
-- From the CLI instead: sqlite3 TL.db "VACUUM;"
--------------------------------------------------------------------------
-- VERIFY
--------------------------------------------------------------------------
-- PRAGMA integrity_check; -- expect: ok
-- SELECT COUNT(*) FROM Notes; -- expect: 1182333, unchanged
-- SELECT COUNT(*) FROM BlogNames; -- expect: 20430
-- SELECT COUNT(*) FROM Blogs WHERE BlogId IS NOT NULL; -- expect: 20418
--
-- Losslessness was proven before this ran, by reconstructing the old text shape
-- from the new schema and diffing it against the original both ways:
-- SELECT COUNT(*) FROM (SELECT * FROM old.Notes EXCEPT SELECT * FROM Rebuilt);
-- SELECT COUNT(*) FROM (SELECT * FROM Rebuilt EXCEPT SELECT * FROM old.Notes);
-- Both returned 0 across all 1,182,333 rows and all 10 columns.