Merge branch 'claude/shrink-tl-db' into master

This commit is contained in:
jim
2026-08-07 20:20:29 -05:00
2 changed files with 211 additions and 15 deletions
+42 -5
View File
@@ -130,13 +130,47 @@ CREATE TABLE "Notes" (
"DatetimeCrawled" TEXT DEFAULT '2/12/26 12am',
"DateModified" TEXT,
"DateCreated" TEXT,
IsActive INTEGER NOT NULL DEFAULT 1,
PRIMARY KEY("RootBlogName","PostID","TimeStamp","Type","NoteBlogName")
);
) WITHOUT ROWID;
CREATE INDEX "Notes_idx_06e01ae3" ON "Notes" ("TimeStamp" DESC);
CREATE INDEX "ix_NoteBlogName01" ON "Notes" ("NoteBlogName");
```
**`WITHOUT ROWID`, since 2026-08-07.** The rows live in the primary key's b-tree
rather than in a rowid table with a separate key index beside it. Nothing about the
SQL surface changes — same columns, same types, same constraint — but two
consequences are worth knowing before adding an index here:
- There is no `rowid` on this table. `SELECT rowid FROM Notes` is an error, and no
code in any of the three apps relied on it.
- A secondary index carries the whole five-column primary key as its row reference
instead of a compact rowid, so indexes on this table are **expensive**.
`ix_NoteBlogName01` costs 58 MB, up from 25 MB before the conversion. It earns
that: Rolodex filters on `NoteBlogName` and the crawler joins on it.
**`Notes_idx_06e01ae3` on `TimeStamp DESC` was dropped at the same time.** It cost
14 MB as a rowid index and would have cost 58 MB after the conversion. It was worth
neither: the crawler's only `TimeStamp` filter (`>= 1535778000`) excludes 786 rows
of 1.18M, Rolodex's default Notes sort carries a three-column tiebreaker that forces
a full sort regardless, and the reply-matching `UPDATE` uses `ABS(TimeStamp - ?) <= 5`,
which no index on `TimeStamp` can serve. The one path that got slower is Rolodex's
Notes page with a date-range filter: 60 ms to 164 ms.
See `../shrink-db.sql` for the full rationale and the applied result.
**`DatetimeCrawled` is `NULL` on 1,148,077 rows, and that is the honest value.** Those
rows previously stored the literal string `'2/12/26 12am'` — this column's own DDL
default, written as a bulk backfill placeholder rather than as a crawl time. They were
set to `NULL` on 2026-08-07, which is what consumers already displayed them as: the
string parses as a date in neither format this schema writes.
Note the trap: **the `DEFAULT '2/12/26 12am'` clause is still in the DDL above.** Any
`INSERT` that omits this column writes the placeholder straight back. The crawler names
it explicitly on every insert, so nothing reintroduces it today, but a new writer that
forgets to would — which is why consumers should keep treating an unparseable value here
as "unknown" rather than assuming `NULL` is now the only such marker.
One row per engagement event. `TimeStamp` is **unix seconds** — unlike every date column
elsewhere in the schema, which are text.
@@ -155,8 +189,11 @@ At 1.19M rows this is the table that dictates how the whole database has to be q
- The only fast access paths are the primary key's leading columns (`RootBlogName`, then
`PostID`) and `ix_NoteBlogName01` on `NoteBlogName`. "Notes received by a blog" and
"notes given by a blog" are both cheap; almost nothing else is.
- Ordering by anything but `TimeStamp` is a full sort of whatever the filters leave.
- `replyText` is `'.'` on 1,174,706 rows — only `reply` notes carry real text.
- **Every** ordering here is a full sort of whatever the filters leave, `TimeStamp`
included. That was already true in practice of the default `TimeStamp` order, whose
tiebreakers forced a sort even while `Notes_idx_06e01ae3` existed; since that index
was dropped on 2026-08-07 it is true unconditionally. Filter first, then sort.
- `replyText` is `'.'` on 1,167,464 rows — only `reply` notes carry real text.
### Referential integrity
@@ -178,7 +215,7 @@ consumer.
| Column | `'.'` rows |
|---|--:|
| `Notes.replyText` | 1,174,706 |
| `Notes.replyText` | 1,167,464 |
| `Posts.Title` | 13,144 |
| `Posts.Body` | 172 |
+159
View File
@@ -0,0 +1,159 @@
-- shrink-db.sql
-- Reduces TL.db from ~267 MB to ~207 MB (-22%) with no application changes,
-- and no visible change in any of the three apps that touch this file.
--
-- The three consumers, and what each one uses:
-- URLNotesGrabberCORE System.Data.SQLite 1.0.119 writes Notes, Posts, Blogs
-- Rolodex (web) Microsoft.Data.Sqlite 10.0 reads all three; soft-deletes via IsActive
-- TumblThree System.Data.SQLite.Core 1.0.119
-- one statement only, ManagerController.cs:972 --
-- "UPDATE Blogs SET IsActive = 0, DateModified = @DateModified
-- WHERE BlogName = @BlogName"
-- Nothing below touches the Blogs table, so TumblThree is
-- unaffected. (Its GlobalDatabaseService talks to TumblThree's
-- own separate FileEntries/BlogFiles database, not this file.)
--
-- WITHOUT ROWID needs SQLite >= 3.8.2 (Dec 2013). All three providers above are
-- 2024-25 builds, an order of magnitude newer, so STEP 2 is readable by all of them.
--
-- Every figure below was measured on a copy of the live 267 MB file, and the
-- result was checked against all three apps' access patterns:
-- PRAGMA integrity_check ....... ok
-- row counts ................... Notes 1182333, Posts 22468, Blogs 188620 (unchanged)
-- Rolodex soft-delete UPDATE ... works
-- Rolodex NoteBlogName filter .. still uses ix_NoteBlogName01
-- crawler INSERT OR IGNORE ..... still dedupes (0 dupes admitted)
-- TumblThree's UPDATE Blogs ..... untouched -- Blogs is not modified by this script
--
-- HOW TO RUN (DB Browser for SQLite):
-- 1. Stop ALL THREE apps: the crawler, the Rolodex web app, and TumblThree.
-- Rolodex holds the file open and checkpoints the WAL, so it must be down,
-- not just idle. TumblThree only opens the file for an instant when you
-- delete a blog, but it can also launch the crawler on its own
-- (UrlNotesGrabberService) -- so close it rather than merely avoiding it.
-- 2. Back up TL.db (copy the 267 MB file somewhere safe).
-- 3. Open TL.db, go to Execute SQL, paste STEP 1-3, run.
-- 4. Click "Write Changes".
-- 5. Run Tools > Compact Database. This is VACUUM; it will not run from the
-- Execute SQL tab because DB Browser keeps a transaction open there.
-- NOTHING SHRINKS ON DISK UNTIL THIS FINISHES.
--
-- Expected: steps 1-3 a couple of minutes, Compact a couple more.
-- Free disk needed during Compact: ~270 MB for the temp copy.
--------------------------------------------------------------------------
-- STEP 1: drop the TimeStamp index (required by STEP 2, not optional)
--------------------------------------------------------------------------
-- Measured cost/benefit:
--
-- * Rolodex's DEFAULT Notes view does not use it. Its sort carries the
-- tiebreaker "RootBlogName, PostID, NoteBlogName", which forces a full sort
-- regardless -- the query plan is byte-identical with and without the index.
-- Sorting.cs:138 already assumes as much, and is right in practice.
-- * The crawler's collect query (DataAccess.cs:1183) filters
-- TimeStamp >= 1535778000, which excludes 786 of 1,182,333 rows (0.07%).
-- A full index scan wearing a disguise. Same measured time without it.
-- * The reply-matching UPDATE uses ABS(TimeStamp - ?) <= 5, which can never
-- use an index on TimeStamp.
-- * It DOES help exactly one path: Rolodex's Notes page with a date-range
-- filter applied. 60 ms -> 164 ms. That is the whole of what is lost.
--
-- And it must go, because after STEP 2 it stops being cheap. A secondary index
-- on a WITHOUT ROWID table carries the full 5-column primary key instead of a
-- compact rowid, so this index grows 14 MB -> 58 MB. Keeping it lands the file
-- at 265 MB instead of 207 MB -- i.e. it cancels the entire exercise to save
-- 100 ms on one filtered view.
DROP INDEX IF EXISTS Notes_idx_06e01ae3;
--------------------------------------------------------------------------
-- STEP 2: rebuild Notes as WITHOUT ROWID (-32 MB)
--------------------------------------------------------------------------
-- Notes has a 5-column composite primary key. In a rowid table SQLite stores
-- that key twice: once in the table, once in sqlite_autoindex_Notes_1 (62 MB).
-- WITHOUT ROWID stores the rows *in* the key's b-tree, so the copy disappears.
--
-- ix_NoteBlogName01 grows 25 -> 58 MB for the reason described above. Net -32 MB.
-- It is kept because Rolodex filters on NoteBlogName and the crawler joins on it.
--
-- Safe: neither codebase references rowid on Notes (grep across both trees,
-- zero matches). IsActive keeps its exact current declaration, which is what
-- Rolodex's ActiveFlag predicate reads.
PRAGMA foreign_keys = off;
CREATE TABLE Notes_new (
"RootBlogName" TEXT,
"PostID" INTEGER,
"NoteBlogName" TEXT,
"TimeStamp" INTEGER,
"Type" TEXT,
"replyText" TEXT DEFAULT '.',
"DatetimeCrawled" TEXT DEFAULT '2/12/26 12am',
"DateModified" TEXT,
"DateCreated" TEXT,
IsActive INTEGER NOT NULL DEFAULT 1,
PRIMARY KEY("RootBlogName","PostID","TimeStamp","Type","NoteBlogName")
) WITHOUT ROWID;
INSERT INTO Notes_new
SELECT RootBlogName, PostID, NoteBlogName, TimeStamp, Type,
replyText, DatetimeCrawled, DateModified, DateCreated, IsActive
FROM Notes;
DROP TABLE Notes;
ALTER TABLE Notes_new RENAME TO Notes;
CREATE INDEX "ix_NoteBlogName01" ON "Notes" ("NoteBlogName");
--------------------------------------------------------------------------
-- STEP 3: clear the DatetimeCrawled placeholder (-13 MB)
--------------------------------------------------------------------------
-- 1,148,077 of 1,182,333 rows hold the literal DDL default '2/12/26 12am' --
-- a backfill placeholder, not a crawl time. SQLite stores all 12 bytes of it
-- on every one of those rows.
--
-- This is UI-NEUTRAL in Rolodex, which is why it is safe despite Rolodex
-- displaying the column. Rolodex reads and sorts it through DateSql.Sortable
-- (DateRange.cs:67), whose CASE matches '____-__-__%' or the 8-character
-- '__/__/__'. The 12-character '2/12/26 12am' matches neither, so Sortable
-- already returns NULL for these rows and the page already renders an em dash
-- and sorts them to the bottom. RolodexRepository.cs:957-963 documents exactly
-- this. Writing a real NULL changes the bytes on disk, not the screen.
--
-- The crawler never reads the column back -- it only writes it on INSERT
-- (DataAccess.cs:713, 721).
UPDATE Notes SET DatetimeCrawled = NULL WHERE DatetimeCrawled = '2/12/26 12am';
--------------------------------------------------------------------------
-- DELIBERATELY NOT DONE: nulling Notes.DateCreated
--------------------------------------------------------------------------
-- An earlier draft of this script also cleared DateCreated = '2026-04-13'
-- (a further -12 MB). Do not. Unlike DatetimeCrawled, that value DOES match
-- Sortable's '____-__-__%' branch, so Rolodex renders it as a real date in the
-- "Created" column on the Notes page and Post detail, and sorts by it. Nulling
-- it would turn visible dates into em dashes and move rows in the sort order.
--------------------------------------------------------------------------
-- STEP 4: Write Changes, then Tools > Compact Database
--------------------------------------------------------------------------
-- Nothing above reclaims disk until VACUUM runs. From the sqlite3 CLI instead:
-- sqlite3 TL.db "VACUUM;"
--------------------------------------------------------------------------
-- VERIFY (run after compacting; file should be ~207 MB)
--------------------------------------------------------------------------
-- PRAGMA integrity_check;
--
-- SELECT 'Notes' t, COUNT(*) n FROM Notes
-- UNION ALL SELECT 'Posts', COUNT(*) FROM Posts
-- UNION ALL SELECT 'Blogs', COUNT(*) FROM Blogs;
-- -- expect 1182333 / 22468 / 188620, unchanged
--
-- SELECT name, SUM(pgsize)/1024/1024 AS mb
-- FROM dbstat GROUP BY name ORDER BY SUM(pgsize) DESC;
-- -- expect Notes 78, ix_NoteBlogName01 58, Posts 50, Blogs 14
--
-- APPLIED 2026-08-07. Actual result: 267.32 MB -> 207.17 MB, integrity_check ok,
-- row counts unchanged, journal_mode still wal. VACUUM took 6 seconds.