diff --git a/URLNotesGrabberCORE/TL.db.md b/URLNotesGrabberCORE/TL.db.md index 8fed093..d89c843 100644 --- a/URLNotesGrabberCORE/TL.db.md +++ b/URLNotesGrabberCORE/TL.db.md @@ -121,22 +121,56 @@ Notable: ```sql CREATE TABLE "Notes" ( - "RootBlogName" TEXT, - "PostID" INTEGER, - "NoteBlogName" TEXT, - "TimeStamp" INTEGER, - "Type" TEXT, - "replyText" TEXT DEFAULT '.', - "DatetimeCrawled" TEXT DEFAULT '2/12/26 12am', - "DateModified" TEXT, - "DateCreated" TEXT, + "RootBlogName" TEXT, + "PostID" INTEGER, + "NoteBlogName" TEXT, + "TimeStamp" INTEGER, + "Type" TEXT, + "replyText" TEXT DEFAULT '.', + "DatetimeCrawled" TEXT DEFAULT '2/12/26 12am', + "DateModified" TEXT, + "DateCreated" TEXT, + IsActive INTEGER NOT NULL DEFAULT 1, PRIMARY KEY("RootBlogName","PostID","TimeStamp","Type","NoteBlogName") -); +) WITHOUT ROWID; -CREATE INDEX "Notes_idx_06e01ae3" ON "Notes" ("TimeStamp" DESC); -CREATE INDEX "ix_NoteBlogName01" ON "Notes" ("NoteBlogName"); +CREATE INDEX "ix_NoteBlogName01" ON "Notes" ("NoteBlogName"); ``` +**`WITHOUT ROWID`, since 2026-08-07.** The rows live in the primary key's b-tree +rather than in a rowid table with a separate key index beside it. Nothing about the +SQL surface changes — same columns, same types, same constraint — but two +consequences are worth knowing before adding an index here: + +- There is no `rowid` on this table. `SELECT rowid FROM Notes` is an error, and no + code in any of the three apps relied on it. +- A secondary index carries the whole five-column primary key as its row reference + instead of a compact rowid, so indexes on this table are **expensive**. + `ix_NoteBlogName01` costs 58 MB, up from 25 MB before the conversion. It earns + that: Rolodex filters on `NoteBlogName` and the crawler joins on it. + +**`Notes_idx_06e01ae3` on `TimeStamp DESC` was dropped at the same time.** It cost +14 MB as a rowid index and would have cost 58 MB after the conversion. It was worth +neither: the crawler's only `TimeStamp` filter (`>= 1535778000`) excludes 786 rows +of 1.18M, Rolodex's default Notes sort carries a three-column tiebreaker that forces +a full sort regardless, and the reply-matching `UPDATE` uses `ABS(TimeStamp - ?) <= 5`, +which no index on `TimeStamp` can serve. The one path that got slower is Rolodex's +Notes page with a date-range filter: 60 ms to 164 ms. + +See `../shrink-db.sql` for the full rationale and the applied result. + +**`DatetimeCrawled` is `NULL` on 1,148,077 rows, and that is the honest value.** Those +rows previously stored the literal string `'2/12/26 12am'` — this column's own DDL +default, written as a bulk backfill placeholder rather than as a crawl time. They were +set to `NULL` on 2026-08-07, which is what consumers already displayed them as: the +string parses as a date in neither format this schema writes. + +Note the trap: **the `DEFAULT '2/12/26 12am'` clause is still in the DDL above.** Any +`INSERT` that omits this column writes the placeholder straight back. The crawler names +it explicitly on every insert, so nothing reintroduces it today, but a new writer that +forgets to would — which is why consumers should keep treating an unparseable value here +as "unknown" rather than assuming `NULL` is now the only such marker. + One row per engagement event. `TimeStamp` is **unix seconds** — unlike every date column elsewhere in the schema, which are text. @@ -155,8 +189,11 @@ At 1.19M rows this is the table that dictates how the whole database has to be q - The only fast access paths are the primary key's leading columns (`RootBlogName`, then `PostID`) and `ix_NoteBlogName01` on `NoteBlogName`. "Notes received by a blog" and "notes given by a blog" are both cheap; almost nothing else is. -- Ordering by anything but `TimeStamp` is a full sort of whatever the filters leave. -- `replyText` is `'.'` on 1,174,706 rows — only `reply` notes carry real text. +- **Every** ordering here is a full sort of whatever the filters leave, `TimeStamp` + included. That was already true in practice of the default `TimeStamp` order, whose + tiebreakers forced a sort even while `Notes_idx_06e01ae3` existed; since that index + was dropped on 2026-08-07 it is true unconditionally. Filter first, then sort. +- `replyText` is `'.'` on 1,167,464 rows — only `reply` notes carry real text. ### Referential integrity @@ -178,7 +215,7 @@ consumer. | Column | `'.'` rows | |---|--:| -| `Notes.replyText` | 1,174,706 | +| `Notes.replyText` | 1,167,464 | | `Posts.Title` | 13,144 | | `Posts.Body` | 172 | diff --git a/shrink-db.sql b/shrink-db.sql new file mode 100644 index 0000000..27179c7 --- /dev/null +++ b/shrink-db.sql @@ -0,0 +1,159 @@ +-- shrink-db.sql +-- Reduces TL.db from ~267 MB to ~207 MB (-22%) with no application changes, +-- and no visible change in any of the three apps that touch this file. +-- +-- The three consumers, and what each one uses: +-- URLNotesGrabberCORE System.Data.SQLite 1.0.119 writes Notes, Posts, Blogs +-- Rolodex (web) Microsoft.Data.Sqlite 10.0 reads all three; soft-deletes via IsActive +-- TumblThree System.Data.SQLite.Core 1.0.119 +-- one statement only, ManagerController.cs:972 -- +-- "UPDATE Blogs SET IsActive = 0, DateModified = @DateModified +-- WHERE BlogName = @BlogName" +-- Nothing below touches the Blogs table, so TumblThree is +-- unaffected. (Its GlobalDatabaseService talks to TumblThree's +-- own separate FileEntries/BlogFiles database, not this file.) +-- +-- WITHOUT ROWID needs SQLite >= 3.8.2 (Dec 2013). All three providers above are +-- 2024-25 builds, an order of magnitude newer, so STEP 2 is readable by all of them. +-- +-- Every figure below was measured on a copy of the live 267 MB file, and the +-- result was checked against all three apps' access patterns: +-- PRAGMA integrity_check ....... ok +-- row counts ................... Notes 1182333, Posts 22468, Blogs 188620 (unchanged) +-- Rolodex soft-delete UPDATE ... works +-- Rolodex NoteBlogName filter .. still uses ix_NoteBlogName01 +-- crawler INSERT OR IGNORE ..... still dedupes (0 dupes admitted) +-- TumblThree's UPDATE Blogs ..... untouched -- Blogs is not modified by this script +-- +-- HOW TO RUN (DB Browser for SQLite): +-- 1. Stop ALL THREE apps: the crawler, the Rolodex web app, and TumblThree. +-- Rolodex holds the file open and checkpoints the WAL, so it must be down, +-- not just idle. TumblThree only opens the file for an instant when you +-- delete a blog, but it can also launch the crawler on its own +-- (UrlNotesGrabberService) -- so close it rather than merely avoiding it. +-- 2. Back up TL.db (copy the 267 MB file somewhere safe). +-- 3. Open TL.db, go to Execute SQL, paste STEP 1-3, run. +-- 4. Click "Write Changes". +-- 5. Run Tools > Compact Database. This is VACUUM; it will not run from the +-- Execute SQL tab because DB Browser keeps a transaction open there. +-- NOTHING SHRINKS ON DISK UNTIL THIS FINISHES. +-- +-- Expected: steps 1-3 a couple of minutes, Compact a couple more. +-- Free disk needed during Compact: ~270 MB for the temp copy. + +-------------------------------------------------------------------------- +-- STEP 1: drop the TimeStamp index (required by STEP 2, not optional) +-------------------------------------------------------------------------- +-- Measured cost/benefit: +-- +-- * Rolodex's DEFAULT Notes view does not use it. Its sort carries the +-- tiebreaker "RootBlogName, PostID, NoteBlogName", which forces a full sort +-- regardless -- the query plan is byte-identical with and without the index. +-- Sorting.cs:138 already assumes as much, and is right in practice. +-- * The crawler's collect query (DataAccess.cs:1183) filters +-- TimeStamp >= 1535778000, which excludes 786 of 1,182,333 rows (0.07%). +-- A full index scan wearing a disguise. Same measured time without it. +-- * The reply-matching UPDATE uses ABS(TimeStamp - ?) <= 5, which can never +-- use an index on TimeStamp. +-- * It DOES help exactly one path: Rolodex's Notes page with a date-range +-- filter applied. 60 ms -> 164 ms. That is the whole of what is lost. +-- +-- And it must go, because after STEP 2 it stops being cheap. A secondary index +-- on a WITHOUT ROWID table carries the full 5-column primary key instead of a +-- compact rowid, so this index grows 14 MB -> 58 MB. Keeping it lands the file +-- at 265 MB instead of 207 MB -- i.e. it cancels the entire exercise to save +-- 100 ms on one filtered view. + +DROP INDEX IF EXISTS Notes_idx_06e01ae3; + +-------------------------------------------------------------------------- +-- STEP 2: rebuild Notes as WITHOUT ROWID (-32 MB) +-------------------------------------------------------------------------- +-- Notes has a 5-column composite primary key. In a rowid table SQLite stores +-- that key twice: once in the table, once in sqlite_autoindex_Notes_1 (62 MB). +-- WITHOUT ROWID stores the rows *in* the key's b-tree, so the copy disappears. +-- +-- ix_NoteBlogName01 grows 25 -> 58 MB for the reason described above. Net -32 MB. +-- It is kept because Rolodex filters on NoteBlogName and the crawler joins on it. +-- +-- Safe: neither codebase references rowid on Notes (grep across both trees, +-- zero matches). IsActive keeps its exact current declaration, which is what +-- Rolodex's ActiveFlag predicate reads. + +PRAGMA foreign_keys = off; + +CREATE TABLE Notes_new ( + "RootBlogName" TEXT, + "PostID" INTEGER, + "NoteBlogName" TEXT, + "TimeStamp" INTEGER, + "Type" TEXT, + "replyText" TEXT DEFAULT '.', + "DatetimeCrawled" TEXT DEFAULT '2/12/26 12am', + "DateModified" TEXT, + "DateCreated" TEXT, + IsActive INTEGER NOT NULL DEFAULT 1, + PRIMARY KEY("RootBlogName","PostID","TimeStamp","Type","NoteBlogName") +) WITHOUT ROWID; + +INSERT INTO Notes_new +SELECT RootBlogName, PostID, NoteBlogName, TimeStamp, Type, + replyText, DatetimeCrawled, DateModified, DateCreated, IsActive + FROM Notes; + +DROP TABLE Notes; +ALTER TABLE Notes_new RENAME TO Notes; + +CREATE INDEX "ix_NoteBlogName01" ON "Notes" ("NoteBlogName"); + +-------------------------------------------------------------------------- +-- STEP 3: clear the DatetimeCrawled placeholder (-13 MB) +-------------------------------------------------------------------------- +-- 1,148,077 of 1,182,333 rows hold the literal DDL default '2/12/26 12am' -- +-- a backfill placeholder, not a crawl time. SQLite stores all 12 bytes of it +-- on every one of those rows. +-- +-- This is UI-NEUTRAL in Rolodex, which is why it is safe despite Rolodex +-- displaying the column. Rolodex reads and sorts it through DateSql.Sortable +-- (DateRange.cs:67), whose CASE matches '____-__-__%' or the 8-character +-- '__/__/__'. The 12-character '2/12/26 12am' matches neither, so Sortable +-- already returns NULL for these rows and the page already renders an em dash +-- and sorts them to the bottom. RolodexRepository.cs:957-963 documents exactly +-- this. Writing a real NULL changes the bytes on disk, not the screen. +-- +-- The crawler never reads the column back -- it only writes it on INSERT +-- (DataAccess.cs:713, 721). + +UPDATE Notes SET DatetimeCrawled = NULL WHERE DatetimeCrawled = '2/12/26 12am'; + +-------------------------------------------------------------------------- +-- DELIBERATELY NOT DONE: nulling Notes.DateCreated +-------------------------------------------------------------------------- +-- An earlier draft of this script also cleared DateCreated = '2026-04-13' +-- (a further -12 MB). Do not. Unlike DatetimeCrawled, that value DOES match +-- Sortable's '____-__-__%' branch, so Rolodex renders it as a real date in the +-- "Created" column on the Notes page and Post detail, and sorts by it. Nulling +-- it would turn visible dates into em dashes and move rows in the sort order. + +-------------------------------------------------------------------------- +-- STEP 4: Write Changes, then Tools > Compact Database +-------------------------------------------------------------------------- +-- Nothing above reclaims disk until VACUUM runs. From the sqlite3 CLI instead: +-- sqlite3 TL.db "VACUUM;" + +-------------------------------------------------------------------------- +-- VERIFY (run after compacting; file should be ~207 MB) +-------------------------------------------------------------------------- +-- PRAGMA integrity_check; +-- +-- SELECT 'Notes' t, COUNT(*) n FROM Notes +-- UNION ALL SELECT 'Posts', COUNT(*) FROM Posts +-- UNION ALL SELECT 'Blogs', COUNT(*) FROM Blogs; +-- -- expect 1182333 / 22468 / 188620, unchanged +-- +-- SELECT name, SUM(pgsize)/1024/1024 AS mb +-- FROM dbstat GROUP BY name ORDER BY SUM(pgsize) DESC; +-- -- expect Notes 78, ix_NoteBlogName01 58, Posts 50, Blogs 14 +-- +-- APPLIED 2026-08-07. Actual result: 267.32 MB -> 207.17 MB, integrity_check ok, +-- row counts unchanged, journal_mode still wal. VACUUM took 6 seconds.