Merge branch 'claude/normalize-notes' into master
This commit is contained in:
+369
-72
@@ -4,12 +4,25 @@ The SQLite database behind **URLNotesGrabberCORE** and its sibling crawlers, and
|
|||||||
[Rolodex](https://git.basso.land/jim/Rolodex) reads.
|
[Rolodex](https://git.basso.land/jim/Rolodex) reads.
|
||||||
|
|
||||||
Everything below was read out of the live file, not inferred from code. Counts are as of
|
Everything below was read out of the live file, not inferred from code. Counts are as of
|
||||||
**2026-07-29**; re-run the queries at the bottom to refresh them.
|
**2026-08-07**; re-run the queries at the bottom to refresh them.
|
||||||
|
|
||||||
- Journal mode: **WAL** — `TL.db-wal` and `TL.db-shm` live beside the file and are part of
|
- Journal mode: **WAL** — `TL.db-wal` and `TL.db-shm` live beside the file and are part of
|
||||||
the database. Copying `TL.db` alone gives you whatever was last checkpointed, not the
|
the database. Copying `TL.db` alone gives you whatever was last checkpointed, not the
|
||||||
current state.
|
current state.
|
||||||
- Page size: 4096.
|
- Page size: 4096. File size: 148 MB.
|
||||||
|
|
||||||
|
> ### ⚠ Breaking change, 2026-08-07: `Notes` holds integer IDs, not names
|
||||||
|
>
|
||||||
|
> `Notes.RootBlogName`, `Notes.NoteBlogName` and `Notes.Type` **no longer exist**. They
|
||||||
|
> are now `RootBlogId`, `NoteBlogId` and `TypeId`, resolved through the new `BlogNames`
|
||||||
|
> and `NoteTypes` tables. Any query naming the old columns fails outright.
|
||||||
|
>
|
||||||
|
> There is no compatibility view. See [porting to the integer
|
||||||
|
> schema](#porting-to-the-integer-schema) for the old-to-new translation of every query
|
||||||
|
> shape the applications use.
|
||||||
|
>
|
||||||
|
> Applied by `../normalize-notes.sql`, which took the file from 207 MB to 148 MB. An
|
||||||
|
> earlier change the same day (`../shrink-db.sql`) took it from 267 MB to 207 MB.
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
@@ -17,13 +30,20 @@ Everything below was read out of the live file, not inferred from code. Counts a
|
|||||||
|
|
||||||
| Table | Rows | What it is |
|
| Table | Rows | What it is |
|
||||||
|---|--:|---|
|
|---|--:|---|
|
||||||
| `Blogs` | 144,367 | The crawl registry — one row per known blog, plus crawl-state flags |
|
| `Blogs` | 188,620 | The crawl registry — one row per known blog, plus crawl-state flags |
|
||||||
| `Posts` | 14,589 | Stored post content. Only 3,602 blogs actually have any |
|
| `Posts` | 22,468 | Stored post content. Only 3,867 blogs actually have any |
|
||||||
| `Notes` | 1,189,604 | The engagement graph: `NoteBlogName` acted on `(RootBlogName, PostID)` |
|
| `Notes` | 1,182,333 | The engagement graph: `NoteBlogId` acted on `(RootBlogId, PostID)` |
|
||||||
|
|
||||||
The engagement graph is the interesting part. 31,888 distinct blogs appear as engagers —
|
…supported by two lookup tables that exist only to keep `Notes` small:
|
||||||
far more than the 3,602 that have stored posts — which is what makes this a social graph
|
|
||||||
rather than a post archive.
|
| Table | Rows | What it is |
|
||||||
|
|---|--:|---|
|
||||||
|
| `BlogNames` | 20,430 | `BlogId` ⇄ `BlogName`. The ID authority for everything in `Notes` |
|
||||||
|
| `NoteTypes` | 5 | `TypeId` ⇄ `Type`. `like`, `reblog`, `reply`, `posted`, `post_attribution` |
|
||||||
|
|
||||||
|
The engagement graph is the interesting part. 20,311 distinct blogs appear as engagers —
|
||||||
|
far more than the 3,867 that have stored posts — which is what makes this a social graph
|
||||||
|
rather than a post archive. Only 2,771 blogs appear as the *root* of a note.
|
||||||
|
|
||||||
### `Blogs`
|
### `Blogs`
|
||||||
|
|
||||||
@@ -42,23 +62,45 @@ CREATE TABLE "Blogs" (
|
|||||||
LikesLastRefreshed INTEGER DEFAULT 0,
|
LikesLastRefreshed INTEGER DEFAULT 0,
|
||||||
LikesLastNewCount INTEGER DEFAULT 0,
|
LikesLastNewCount INTEGER DEFAULT 0,
|
||||||
TTFolderPath TEXT,
|
TTFolderPath TEXT,
|
||||||
|
BlogId INTEGER,
|
||||||
PRIMARY KEY("BlogName")
|
PRIMARY KEY("BlogName")
|
||||||
);
|
);
|
||||||
|
|
||||||
|
CREATE INDEX ix_Blogs_BlogId ON Blogs (BlogId);
|
||||||
```
|
```
|
||||||
|
|
||||||
`BlogName` is the primary key, so it is the only indexed way in. There is no index on any
|
`BlogName` is the primary key, so it is the only indexed way in by name. There is no index
|
||||||
flag or date — filtering or sorting on those scans all 144k rows, which is affordable
|
on any flag or date — filtering or sorting on those scans all 188k rows, which is
|
||||||
here and is not on `Notes`.
|
affordable here and is not on `Notes`.
|
||||||
|
|
||||||
Flag distribution: `IsActive = 1` on 144,366 of 144,367 rows, `HasBeenOutput = 1` on
|
**`BlogId` is new as of 2026-08-07 and is the join key to `Notes`.** It exists so that
|
||||||
5,369, `ByLikes = 1` on 2. `IsActive` carries a second meaning as of Rolodex — see
|
`Notes` can reach `Blogs` in a single integer hop rather than going through `BlogNames`
|
||||||
|
and ending in a text comparison:
|
||||||
|
|
||||||
|
```sql
|
||||||
|
-- what you want
|
||||||
|
FROM Blogs B JOIN Notes N ON N.NoteBlogId = B.BlogId
|
||||||
|
|
||||||
|
-- not this
|
||||||
|
FROM Blogs B JOIN BlogNames BN ON BN.BlogName = B.BlogName
|
||||||
|
JOIN Notes N ON N.NoteBlogId = BN.BlogId
|
||||||
|
```
|
||||||
|
|
||||||
|
**`BlogId` is NULL on 168,202 of 188,620 rows** — every blog that has never appeared in a
|
||||||
|
note. That is the large majority, and it is not an error: the registry is far bigger than
|
||||||
|
the engagement graph. An inner join on `BlogId` therefore silently drops those blogs,
|
||||||
|
which is usually what you want for engagement queries and is wrong for registry listings.
|
||||||
|
|
||||||
|
Flag distribution: `IsActive = 1` on 188,601 of 188,620 rows, `HasBeenOutput = 1` on
|
||||||
|
5,059, `ByLikes = 1` on 2. `IsActive` carries a second meaning as of Rolodex — see
|
||||||
[`Blogs.IsActive`](#blogsisactive--now-written-by-two-applications) below.
|
[`Blogs.IsActive`](#blogsisactive--now-written-by-two-applications) below.
|
||||||
|
|
||||||
The columns after `DateCreated` were added later by `ALTER TABLE`, which is why they carry
|
The columns after `DateCreated` were added later by `ALTER TABLE`, which is why they carry
|
||||||
no quoting in the stored DDL. That is the normal way this schema grows.
|
no quoting in the stored DDL. That is the normal way this schema grows, and `BlogId` is
|
||||||
|
the newest example.
|
||||||
|
|
||||||
**`DateAdded` is not written consistently.** 126,423 rows hold ISO `yyyy-MM-dd HH:mm:ss`;
|
**`DateAdded` is not written consistently.** 170,677 rows hold ISO `yyyy-MM-dd HH:mm:ss`;
|
||||||
17,944 hold US-format `M/d/yy` from a bulk import. As text those two sort into different
|
17,943 hold US-format `M/d/yy` from a bulk import. As text those two sort into different
|
||||||
parts of the table, so anything ordering or range-filtering on this column has to
|
parts of the table, so anything ordering or range-filtering on this column has to
|
||||||
normalise first — see `DateSql` in Rolodex.
|
normalise first — see `DateSql` in Rolodex.
|
||||||
|
|
||||||
@@ -100,54 +142,73 @@ CREATE TABLE "Posts" (
|
|||||||
);
|
);
|
||||||
```
|
```
|
||||||
|
|
||||||
**The key is `(BlogName, PostID)`, not `PostID`.** This matters more than it looks: 325
|
**The key is `(BlogName, PostID)`, not `PostID`.** This matters more than it looks: 345
|
||||||
post IDs exist under more than one blog, so an ID on its own is both ambiguous *and*
|
post IDs exist under more than one blog, so an ID on its own is both ambiguous *and*
|
||||||
unindexed. Any lookup should carry the blog name, and a batch lookup should group by blog
|
unindexed. Any lookup should carry the blog name, and a batch lookup should group by blog
|
||||||
so it stays on the leading column of the key.
|
so it stays on the leading column of the key.
|
||||||
|
|
||||||
Notable:
|
Notable:
|
||||||
|
|
||||||
- **`PostType` is `NULL` on all 14,589 rows.** The column exists but nothing has ever
|
- **`PostType` is now mostly populated: 20,679 of 22,468 rows, leaving 1,789 `NULL`.**
|
||||||
populated it. Treat it as unpopulated rather than as a type discriminator.
|
This reverses what earlier revisions of this document said — the column really was empty
|
||||||
- `HasImage = 1` on 14,268 rows — nearly all of them. It records that the post *had* a
|
on every row, and something has since started writing it. Anything that treated it as
|
||||||
picture, not that a usable URL was kept, so it is not a reliable predictor that anything
|
permanently unset, or derived the type from post content instead, should be re-examined
|
||||||
will render.
|
against the live data. Rolodex still derives it.
|
||||||
|
- `HasImage = 1` on 14,026 rows. It records that the post *had* a picture, not that a
|
||||||
|
usable URL was kept, so it is not a reliable predictor that anything will render.
|
||||||
- `PhotoURL` is largely unused; in practice the image markup lives inside `Body`.
|
- `PhotoURL` is largely unused; in practice the image markup lives inside `Body`.
|
||||||
- `NotFound = 1` on 4,663 rows — posts that have since been deleted upstream.
|
- `NotFound = 1` on 4,712 rows — posts that have since been deleted upstream.
|
||||||
- The content columns (`Body`, `Quote`, `Question`, `Answer`, …) are the heavy ones. List
|
- The content columns (`Body`, `Quote`, `Question`, `Answer`, …) are the heavy ones. List
|
||||||
views should not select them.
|
views should not select them.
|
||||||
|
|
||||||
### `Notes`
|
### `Notes`
|
||||||
|
|
||||||
```sql
|
```sql
|
||||||
CREATE TABLE "Notes" (
|
CREATE TABLE Notes (
|
||||||
"RootBlogName" TEXT,
|
RootBlogId INTEGER NOT NULL,
|
||||||
"PostID" INTEGER,
|
PostID INTEGER NOT NULL,
|
||||||
"NoteBlogName" TEXT,
|
NoteBlogId INTEGER NOT NULL,
|
||||||
"TimeStamp" INTEGER,
|
TimeStamp INTEGER NOT NULL,
|
||||||
"Type" TEXT,
|
TypeId INTEGER NOT NULL,
|
||||||
"replyText" TEXT DEFAULT '.',
|
replyText TEXT,
|
||||||
"DatetimeCrawled" TEXT DEFAULT '2/12/26 12am',
|
DatetimeCrawled TEXT,
|
||||||
"DateModified" TEXT,
|
DateModified TEXT,
|
||||||
"DateCreated" TEXT,
|
DateCreated TEXT,
|
||||||
IsActive INTEGER NOT NULL DEFAULT 1,
|
IsActive INTEGER NOT NULL DEFAULT 1,
|
||||||
PRIMARY KEY("RootBlogName","PostID","TimeStamp","Type","NoteBlogName")
|
PRIMARY KEY (RootBlogId, PostID, TimeStamp, TypeId, NoteBlogId)
|
||||||
) WITHOUT ROWID;
|
) WITHOUT ROWID;
|
||||||
|
|
||||||
CREATE INDEX "ix_NoteBlogName01" ON "Notes" ("NoteBlogName");
|
CREATE INDEX ix_Notes_NoteBlogId ON Notes (NoteBlogId);
|
||||||
```
|
```
|
||||||
|
|
||||||
**`WITHOUT ROWID`, since 2026-08-07.** The rows live in the primary key's b-tree
|
**Integer IDs since 2026-08-07 — this is the breaking change.** `RootBlogName`,
|
||||||
rather than in a rowid table with a separate key index beside it. Nothing about the
|
`NoteBlogName` and `Type` are gone, replaced by `RootBlogId`, `NoteBlogId` and `TypeId`.
|
||||||
SQL surface changes — same columns, same types, same constraint — but two
|
Resolve them through [`BlogNames`](#blognames) and [`NoteTypes`](#notetypes), or join
|
||||||
consequences are worth knowing before adding an index here:
|
straight to `Blogs` on `BlogId`. The old names were text repeated on 1.18 million rows,
|
||||||
|
in the table *and* in every index over it; the swap took the file from 207 MB to 148 MB.
|
||||||
|
|
||||||
|
The **primary key column order is deliberately unchanged**, so the leading-prefix access
|
||||||
|
patterns callers already depend on still hold: `(RootBlogId)` and `(RootBlogId, PostID)`
|
||||||
|
remain cheap prefixes, exactly as `(RootBlogName)` and `(RootBlogName, PostID)` were.
|
||||||
|
|
||||||
|
Two nulls-and-defaults differences from the old DDL, both intentional:
|
||||||
|
|
||||||
|
- `replyText` and `DatetimeCrawled` **no longer carry column defaults**. The old table
|
||||||
|
defaulted them to `'.'` and `'2/12/26 12am'`, which is how 1.1M rows acquired
|
||||||
|
placeholder values nobody wrote. New rows now get `NULL` unless a writer supplies
|
||||||
|
something. The crawler names both columns explicitly, so its behaviour is unchanged.
|
||||||
|
- The five key columns are now `NOT NULL`. They always were in practice.
|
||||||
|
|
||||||
|
**`WITHOUT ROWID`, since earlier the same day.** The rows live in the primary key's
|
||||||
|
b-tree rather than in a rowid table with a separate key index beside it. Two consequences
|
||||||
|
matter before adding an index here:
|
||||||
|
|
||||||
- There is no `rowid` on this table. `SELECT rowid FROM Notes` is an error, and no
|
- There is no `rowid` on this table. `SELECT rowid FROM Notes` is an error, and no
|
||||||
code in any of the three apps relied on it.
|
code in any of the three apps relied on it.
|
||||||
- A secondary index carries the whole five-column primary key as its row reference
|
- A secondary index carries the whole five-column primary key as its row reference
|
||||||
instead of a compact rowid, so indexes on this table are **expensive**.
|
instead of a compact rowid, so indexes here are **expensive** — though far less so
|
||||||
`ix_NoteBlogName01` costs 58 MB, up from 25 MB before the conversion. It earns
|
than before, now that the key is five integers rather than three integers and two
|
||||||
that: Rolodex filters on `NoteBlogName` and the crawler joins on it.
|
strings. `ix_Notes_NoteBlogId` costs 27 MB; its text predecessor cost 58 MB.
|
||||||
|
|
||||||
**`Notes_idx_06e01ae3` on `TimeStamp DESC` was dropped at the same time.** It cost
|
**`Notes_idx_06e01ae3` on `TimeStamp DESC` was dropped at the same time.** It cost
|
||||||
14 MB as a rowid index and would have cost 58 MB after the conversion. It was worth
|
14 MB as a rowid index and would have cost 58 MB after the conversion. It was worth
|
||||||
@@ -174,37 +235,247 @@ as "unknown" rather than assuming `NULL` is now the only such marker.
|
|||||||
One row per engagement event. `TimeStamp` is **unix seconds** — unlike every date column
|
One row per engagement event. `TimeStamp` is **unix seconds** — unlike every date column
|
||||||
elsewhere in the schema, which are text.
|
elsewhere in the schema, which are text.
|
||||||
|
|
||||||
| `Type` | Rows | Share |
|
At 1.18M rows this is the table that dictates how the whole database has to be queried:
|
||||||
|---|--:|--:|
|
|
||||||
| `like` | 947,955 | 79.7% |
|
|
||||||
| `reblog` | 224,323 | 18.9% |
|
|
||||||
| `reply` | 15,201 | 1.3% |
|
|
||||||
| `posted` | 2,106 | 0.2% |
|
|
||||||
| `post_attribution` | 19 | — |
|
|
||||||
|
|
||||||
At 1.19M rows this is the table that dictates how the whole database has to be queried:
|
|
||||||
|
|
||||||
- **Nothing should run an unbounded `SELECT` or a bare `COUNT(*)` here.** A count scans
|
- **Nothing should run an unbounded `SELECT` or a bare `COUNT(*)` here.** A count scans
|
||||||
the lot on every call.
|
the lot on every call.
|
||||||
- The only fast access paths are the primary key's leading columns (`RootBlogName`, then
|
- The only fast access paths are the primary key's leading columns (`RootBlogId`, then
|
||||||
`PostID`) and `ix_NoteBlogName01` on `NoteBlogName`. "Notes received by a blog" and
|
`PostID`) and `ix_Notes_NoteBlogId` on `NoteBlogId`. "Notes received by a blog" and
|
||||||
"notes given by a blog" are both cheap; almost nothing else is.
|
"notes given by a blog" are both cheap; almost nothing else is.
|
||||||
- **Every** ordering here is a full sort of whatever the filters leave, `TimeStamp`
|
- **Every** ordering here is a full sort of whatever the filters leave, `TimeStamp`
|
||||||
included. That was already true in practice of the default `TimeStamp` order, whose
|
included. Filter first, then sort.
|
||||||
tiebreakers forced a sort even while `Notes_idx_06e01ae3` existed; since that index
|
- `replyText` is `'.'` on 1,167,464 rows — only `reply` notes carry real text. Those
|
||||||
was dropped on 2026-08-07 it is true unconditionally. Filter first, then sort.
|
dots are inherited from the old column default; new rows get `NULL` instead.
|
||||||
- `replyText` is `'.'` on 1,167,464 rows — only `reply` notes carry real text.
|
|
||||||
|
**Resolve IDs by filtering the lookup, not by scanning `Notes`.** The lookup tables are
|
||||||
|
tiny and uniquely indexed, so pushing a name predicate into them costs nothing and lets
|
||||||
|
the `Notes` index do the work:
|
||||||
|
|
||||||
|
```sql
|
||||||
|
-- good: BlogNames resolves the name, then the index is searched
|
||||||
|
SELECT * FROM Notes
|
||||||
|
WHERE NoteBlogId = (SELECT BlogId FROM BlogNames WHERE BlogName = ?);
|
||||||
|
|
||||||
|
-- also good, same plan
|
||||||
|
SELECT n.* FROM Notes n
|
||||||
|
JOIN BlogNames b ON b.BlogId = n.NoteBlogId
|
||||||
|
WHERE b.BlogName = ?;
|
||||||
|
```
|
||||||
|
|
||||||
|
### `BlogNames`
|
||||||
|
|
||||||
|
```sql
|
||||||
|
CREATE TABLE BlogNames (
|
||||||
|
BlogId INTEGER PRIMARY KEY,
|
||||||
|
BlogName TEXT NOT NULL UNIQUE
|
||||||
|
);
|
||||||
|
```
|
||||||
|
|
||||||
|
20,430 rows — every name appearing in `Notes` as either participant, and nothing else.
|
||||||
|
This is the **ID authority**: `Notes.RootBlogId` and `Notes.NoteBlogId` both point here,
|
||||||
|
and `Blogs.BlogId` is a copy of the value for the blogs that have one.
|
||||||
|
|
||||||
|
**12 of these names have no `Blogs` row.** The registry has never been a superset of the
|
||||||
|
engagement graph and still is not, so resolving an ID through `Blogs` rather than
|
||||||
|
`BlogNames` will occasionally find nothing. Use `BlogNames` when you need the name itself
|
||||||
|
and `Blogs` when you need registry columns.
|
||||||
|
|
||||||
|
IDs are assigned by SQLite and are **stable**: they are stored in over a million `Notes`
|
||||||
|
rows. Never renumber them. A blog that is renamed upstream should get a new row, not an
|
||||||
|
edit to an existing one, unless every `Notes` reference is migrated with it.
|
||||||
|
|
||||||
|
### `NoteTypes`
|
||||||
|
|
||||||
|
```sql
|
||||||
|
CREATE TABLE NoteTypes (
|
||||||
|
TypeId INTEGER PRIMARY KEY,
|
||||||
|
Type TEXT NOT NULL UNIQUE
|
||||||
|
);
|
||||||
|
```
|
||||||
|
|
||||||
|
| `TypeId` | `Type` | Rows | Share |
|
||||||
|
|--:|---|--:|--:|
|
||||||
|
| 1 | `like` | 945,167 | 79.9% |
|
||||||
|
| 2 | `reblog` | 219,203 | 18.5% |
|
||||||
|
| 3 | `reply` | 15,345 | 1.3% |
|
||||||
|
| 4 | `posted` | 2,617 | 0.2% |
|
||||||
|
| 5 | `post_attribution` | 1 | — |
|
||||||
|
|
||||||
|
The set is fixed in practice, but it is a table rather than a `CHECK` constraint so that
|
||||||
|
adding a type is an `INSERT` and not a schema migration. **The IDs above are stored in
|
||||||
|
`Notes` and must not be reassigned.**
|
||||||
|
|
||||||
|
Five rows means the lookup is effectively free; write `t.Type = 'reblog'` and let SQLite
|
||||||
|
resolve it, or hardcode the ID if you prefer — both are fine, but hardcoding ties your
|
||||||
|
code to this table's contents, so prefer the join in anything long-lived.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Porting to the integer schema
|
||||||
|
|
||||||
|
Everything here was checked against the live 148 MB file. There are roughly 14 affected
|
||||||
|
call sites in `DataAccess.cs` and 16 in `RolodexRepository.cs`. TumblThree needs no
|
||||||
|
changes — its single statement touches `Blogs.IsActive` and `BlogName` only.
|
||||||
|
|
||||||
|
### Column mapping
|
||||||
|
|
||||||
|
| Was | Is now | Resolve via |
|
||||||
|
|---|---|---|
|
||||||
|
| `Notes.RootBlogName` | `Notes.RootBlogId` | `BlogNames.BlogId` → `.BlogName` |
|
||||||
|
| `Notes.NoteBlogName` | `Notes.NoteBlogId` | `BlogNames.BlogId` → `.BlogName` |
|
||||||
|
| `Notes.Type` | `Notes.TypeId` | `NoteTypes.TypeId` → `.Type` |
|
||||||
|
| `ix_NoteBlogName01` | `ix_Notes_NoteBlogId` | — |
|
||||||
|
|
||||||
|
`PostID`, `TimeStamp`, `replyText`, `DatetimeCrawled`, `DateModified`, `DateCreated` and
|
||||||
|
`IsActive` are unchanged.
|
||||||
|
|
||||||
|
### Filtering by a blog name
|
||||||
|
|
||||||
|
```sql
|
||||||
|
-- was
|
||||||
|
WHERE NoteBlogName = @Name
|
||||||
|
|
||||||
|
-- now, either form; both search ix_Notes_NoteBlogId after a unique-index lookup
|
||||||
|
WHERE NoteBlogId = (SELECT BlogId FROM BlogNames WHERE BlogName = @Name)
|
||||||
|
-- or
|
||||||
|
JOIN BlogNames b ON b.BlogId = n.NoteBlogId WHERE b.BlogName = @Name
|
||||||
|
```
|
||||||
|
|
||||||
|
Measured 73 ms against 63 ms for the old text form on the busiest blog — the extra hop is
|
||||||
|
a unique-index probe on a 20k-row table and does not show.
|
||||||
|
|
||||||
|
### Joining `Notes` to `Blogs`
|
||||||
|
|
||||||
|
This is the join to get right; it is the most common shape in both applications.
|
||||||
|
|
||||||
|
```sql
|
||||||
|
-- was
|
||||||
|
FROM Blogs B INNER JOIN Notes N ON N.NoteBlogName = B.BlogName
|
||||||
|
|
||||||
|
-- now: one integer hop, using the new Blogs.BlogId
|
||||||
|
FROM Blogs B INNER JOIN Notes N ON N.NoteBlogId = B.BlogId
|
||||||
|
```
|
||||||
|
|
||||||
|
Do **not** route this through `BlogNames` — that adds a hop and ends in the text
|
||||||
|
comparison the change was meant to remove.
|
||||||
|
|
||||||
|
### Selecting a name back out
|
||||||
|
|
||||||
|
```sql
|
||||||
|
-- was
|
||||||
|
SELECT NoteBlogName AS blogName, COUNT(*) FROM Notes ... GROUP BY NoteBlogName
|
||||||
|
|
||||||
|
-- now
|
||||||
|
SELECT bn.BlogName AS blogName, COUNT(*)
|
||||||
|
FROM Notes n JOIN BlogNames bn ON bn.BlogId = n.NoteBlogId
|
||||||
|
... GROUP BY bn.BlogName
|
||||||
|
```
|
||||||
|
|
||||||
|
Group by `n.NoteBlogId` instead of `bn.BlogName` when you only need the name for display —
|
||||||
|
grouping on the integer is cheaper and the name comes along for free.
|
||||||
|
|
||||||
|
### Filtering by type
|
||||||
|
|
||||||
|
```sql
|
||||||
|
-- was
|
||||||
|
WHERE type IN ('reblog', 'reply', 'posted')
|
||||||
|
|
||||||
|
-- now
|
||||||
|
WHERE TypeId IN (SELECT TypeId FROM NoteTypes WHERE Type IN ('reblog','reply','posted'))
|
||||||
|
-- or, equivalently
|
||||||
|
JOIN NoteTypes t ON t.TypeId = n.TypeId WHERE t.Type IN ('reblog','reply','posted')
|
||||||
|
```
|
||||||
|
|
||||||
|
`WHERE TypeId IN (2,3,4)` also works and is marginally faster, but hardcodes this table's
|
||||||
|
contents into application code. Prefer the lookup outside of hot paths.
|
||||||
|
|
||||||
|
Note the negated form needs care: `type NOT IN ('reblog','reply','posted')` becomes
|
||||||
|
`TypeId NOT IN (SELECT TypeId FROM NoteTypes WHERE Type IN (...))`, which is correct only
|
||||||
|
because `TypeId` is `NOT NULL`.
|
||||||
|
|
||||||
|
### Inserting a note
|
||||||
|
|
||||||
|
The crawler must ensure both names have IDs first. `INSERT OR IGNORE` on `BlogNames` is
|
||||||
|
the whole of it — no read-back, no round trip, safe to run every time:
|
||||||
|
|
||||||
|
```sql
|
||||||
|
INSERT OR IGNORE INTO BlogNames (BlogName) VALUES (@rootBlogName);
|
||||||
|
INSERT OR IGNORE INTO BlogNames (BlogName) VALUES (@noteBlogName);
|
||||||
|
|
||||||
|
INSERT OR IGNORE INTO Notes
|
||||||
|
(RootBlogId, PostID, NoteBlogId, TimeStamp, TypeId,
|
||||||
|
DatetimeCrawled, DateModified, DateCreated)
|
||||||
|
SELECT (SELECT BlogId FROM BlogNames WHERE BlogName = @rootBlogName),
|
||||||
|
@PostID,
|
||||||
|
(SELECT BlogId FROM BlogNames WHERE BlogName = @noteBlogName),
|
||||||
|
@TimeStamp,
|
||||||
|
(SELECT TypeId FROM NoteTypes WHERE Type = @Type),
|
||||||
|
@DatetimeCrawled, @DateModified, @DateCreated;
|
||||||
|
```
|
||||||
|
|
||||||
|
Verified: a genuinely new note inserts, and re-running the identical statement inserts 0.
|
||||||
|
Run all three statements in one transaction so a crash cannot leave a name registered
|
||||||
|
with no note.
|
||||||
|
|
||||||
|
**The duplicate-key error message has changed.** `DataAccess.cs` compares against the
|
||||||
|
literal string
|
||||||
|
|
||||||
|
```
|
||||||
|
UNIQUE constraint failed: Notes.RootBlogName, Notes.PostID, Notes.TimeStamp, Notes.Type, Notes.NoteBlogName
|
||||||
|
```
|
||||||
|
|
||||||
|
at two call sites to decide whether to swallow an exception. SQLite now emits the *new*
|
||||||
|
column names, so those comparisons no longer match and real errors will surface where
|
||||||
|
they used to be silently ignored — or vice versa. Both sites need updating.
|
||||||
|
|
||||||
|
### Updating notes
|
||||||
|
|
||||||
|
Predicates translate the same way. The reply-matching update, which cannot use an index
|
||||||
|
on `TimeStamp` either before or after:
|
||||||
|
|
||||||
|
```sql
|
||||||
|
-- now
|
||||||
|
UPDATE Notes SET replyText = @replyText, DateModified = @dateModified
|
||||||
|
WHERE NoteBlogId = (SELECT BlogId FROM BlogNames WHERE BlogName = @noteBlogName)
|
||||||
|
AND ABS(TimeStamp - @TimeStamp) <= 5
|
||||||
|
AND TypeId = (SELECT TypeId FROM NoteTypes WHERE Type = 'reply')
|
||||||
|
AND (replyText IS NULL OR replyText = '' OR replyText = '.')
|
||||||
|
AND (replyText IS NULL OR replyText <> @replyText);
|
||||||
|
```
|
||||||
|
|
||||||
|
Rolodex's soft-delete updates need no change beyond the `WHERE` clause — they set
|
||||||
|
`IsActive`, which is untouched.
|
||||||
|
|
||||||
|
### Three traps
|
||||||
|
|
||||||
|
**`Blogs.BlogId` is NULL on 168,202 of 188,620 rows.** Any inner join on it silently drops
|
||||||
|
every blog that has never appeared in a note. Correct for engagement queries; wrong for
|
||||||
|
registry listings, which need a `LEFT JOIN` or no join at all.
|
||||||
|
|
||||||
|
**12 names in `BlogNames` have no `Blogs` row.** Resolving an ID to a name through `Blogs`
|
||||||
|
will occasionally find nothing. Use `BlogNames` for names and `Blogs` for registry columns.
|
||||||
|
|
||||||
|
**IDs are stable and must stay so.** `BlogNames.BlogId` and `NoteTypes.TypeId` are stored
|
||||||
|
in over a million `Notes` rows. Never renumber. A blog renamed upstream gets a new row,
|
||||||
|
not an edited one, unless every `Notes` reference migrates with it.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
### Referential integrity
|
### Referential integrity
|
||||||
|
|
||||||
There are no foreign keys, and the tables do not perfectly agree:
|
There are no foreign keys, and the tables do not perfectly agree:
|
||||||
|
|
||||||
- 4 `Posts` rows name a blog with no `Blogs` row.
|
- 4 `Posts` rows name a blog with no `Blogs` row.
|
||||||
- 15 of the 31,888 distinct engagers have no `Blogs` row.
|
- 12 of the 20,430 names in `BlogNames` have no `Blogs` row.
|
||||||
|
|
||||||
So a name appearing in `Notes` or `Posts` is not a guarantee that the registry knows about
|
So a name appearing in `Notes` or `Posts` is not a guarantee that the registry knows about
|
||||||
it. Joins from those tables back to `Blogs` should tolerate a miss.
|
it. Joins from those tables back to `Blogs` should tolerate a miss.
|
||||||
|
|
||||||
|
The integer schema does not fix this and was not meant to. `BlogNames` is deliberately
|
||||||
|
built from `Notes` rather than from `Blogs`, precisely so that the 12 unregistered
|
||||||
|
engagers keep their IDs and their rows. Had it been built from the registry, those notes
|
||||||
|
would have been dropped by the migration's inner joins.
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
## The `'.'` placeholder convention
|
## The `'.'` placeholder convention
|
||||||
@@ -216,9 +487,13 @@ consumer.
|
|||||||
| Column | `'.'` rows |
|
| Column | `'.'` rows |
|
||||||
|---|--:|
|
|---|--:|
|
||||||
| `Notes.replyText` | 1,167,464 |
|
| `Notes.replyText` | 1,167,464 |
|
||||||
| `Posts.Title` | 13,144 |
|
| `Posts.Title` | 12,562 |
|
||||||
| `Posts.Body` | 172 |
|
| `Posts.Body` | 172 |
|
||||||
|
|
||||||
|
`Notes.replyText` and `Notes.DatetimeCrawled` **no longer carry column defaults** as of
|
||||||
|
the integer migration, so new note rows get `NULL` rather than a placeholder. The dots
|
||||||
|
already in `replyText` were not rewritten — cleaning is still required on read.
|
||||||
|
|
||||||
Any query whose output reaches a human should collapse it:
|
Any query whose output reaches a human should collapse it:
|
||||||
|
|
||||||
```sql
|
```sql
|
||||||
@@ -251,9 +526,12 @@ Crawler bookkeeping. Rolodex ignores all of these.
|
|||||||
`DataAccess.cs` joins on it to decide what to collect:
|
`DataAccess.cs` joins on it to decide what to collect:
|
||||||
|
|
||||||
```sql
|
```sql
|
||||||
SELECT NoteBlogName, count(*) FROM notes
|
-- as it will read after the integer-schema port; see the porting guide above
|
||||||
INNER JOIN blogs ON blogs.BlogName = notes.NoteBlogName
|
SELECT bn.BlogName, count(*)
|
||||||
WHERE blogs.IsActive = @isActive AND ...
|
FROM Notes n
|
||||||
|
JOIN Blogs b ON b.BlogId = n.NoteBlogId
|
||||||
|
JOIN BlogNames bn ON bn.BlogId = n.NoteBlogId
|
||||||
|
WHERE b.IsActive = @isActive AND ...
|
||||||
```
|
```
|
||||||
|
|
||||||
Nothing inside the crawler *writes* it — it is an input, set from outside.
|
Nothing inside the crawler *writes* it — it is an input, set from outside.
|
||||||
@@ -299,12 +577,16 @@ under its Posts and Notes pages. Removing a blog hides the blog, not what it col
|
|||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
## `Posts.IsActive` and `Notes.IsActive` — optional, and not in this database yet
|
## `Posts.IsActive` and `Notes.IsActive` — present, and written from outside
|
||||||
|
|
||||||
The same flag is being extended to the two content tables, with the same meaning: `0` is
|
The same flag extends to the two content tables, with the same meaning: `0` is removed,
|
||||||
removed, anything else — including `NULL` — is live. **Neither column exists in the live
|
anything else — including `NULL` — is live. **Both columns now exist in the live `TL.db`**
|
||||||
`TL.db` as of 2026-07-29**; the DDL quoted above for `Posts` and `Notes` is complete. Like
|
and are included in the DDL quoted above. As of 2026-08-07, `Posts.IsActive = 0` on 5,900
|
||||||
`Blogs.IsActive`, they are written from outside this crawler.
|
rows and `Notes.IsActive = 0` on none. Like `Blogs.IsActive`, they are written from
|
||||||
|
outside this crawler.
|
||||||
|
|
||||||
|
On `Notes` the column is `INTEGER NOT NULL DEFAULT 1`, so a `NULL` cannot occur there;
|
||||||
|
`Posts` and `Blogs` are laxer, which is why the predicate below still uses `COALESCE`.
|
||||||
|
|
||||||
The crawler therefore treats both as optional, and as nothing it owns:
|
The crawler therefore treats both as optional, and as nothing it owns:
|
||||||
|
|
||||||
@@ -344,11 +626,19 @@ handled:
|
|||||||
|
|
||||||
```sql
|
```sql
|
||||||
SELECT 'Blogs', COUNT(*) FROM Blogs
|
SELECT 'Blogs', COUNT(*) FROM Blogs
|
||||||
UNION ALL SELECT 'Posts', COUNT(*) FROM Posts
|
UNION ALL SELECT 'Posts', COUNT(*) FROM Posts
|
||||||
UNION ALL SELECT 'Notes', COUNT(*) FROM Notes;
|
UNION ALL SELECT 'Notes', COUNT(*) FROM Notes
|
||||||
|
UNION ALL SELECT 'BlogNames', COUNT(*) FROM BlogNames;
|
||||||
|
|
||||||
-- note type mix
|
-- note type mix (joins NoteTypes; Notes.Type no longer exists)
|
||||||
SELECT Type, COUNT(*) FROM Notes GROUP BY Type ORDER BY 2 DESC;
|
SELECT t.Type, COUNT(*)
|
||||||
|
FROM Notes n JOIN NoteTypes t ON t.TypeId = n.TypeId
|
||||||
|
GROUP BY t.Type ORDER BY 2 DESC;
|
||||||
|
|
||||||
|
-- how much of the registry participates in the engagement graph
|
||||||
|
SELECT COUNT(*) FILTER (WHERE BlogId IS NOT NULL) AS with_notes,
|
||||||
|
COUNT(*) FILTER (WHERE BlogId IS NULL) AS without_notes
|
||||||
|
FROM Blogs;
|
||||||
|
|
||||||
-- the two date shapes in Blogs.DateAdded
|
-- the two date shapes in Blogs.DateAdded
|
||||||
SELECT CASE WHEN DateAdded LIKE '____-__-__%' THEN 'ISO' ELSE 'US' END, COUNT(*)
|
SELECT CASE WHEN DateAdded LIKE '____-__-__%' THEN 'ISO' ELSE 'US' END, COUNT(*)
|
||||||
@@ -361,6 +651,13 @@ SELECT COUNT(*) FROM (
|
|||||||
-- rows that reference a blog the registry does not have
|
-- rows that reference a blog the registry does not have
|
||||||
SELECT COUNT(*) FROM Posts p
|
SELECT COUNT(*) FROM Posts p
|
||||||
WHERE NOT EXISTS (SELECT 1 FROM Blogs b WHERE b.BlogName = p.BlogName);
|
WHERE NOT EXISTS (SELECT 1 FROM Blogs b WHERE b.BlogName = p.BlogName);
|
||||||
|
|
||||||
|
SELECT COUNT(*) FROM BlogNames bn
|
||||||
|
WHERE NOT EXISTS (SELECT 1 FROM Blogs b WHERE b.BlogName = bn.BlogName);
|
||||||
|
|
||||||
|
-- space by object, to see where the file actually goes
|
||||||
|
SELECT name, SUM(pgsize)/1024/1024 AS mb
|
||||||
|
FROM dbstat GROUP BY name ORDER BY SUM(pgsize) DESC;
|
||||||
```
|
```
|
||||||
|
|
||||||
Open the file read-only so an inspection can never disturb a running crawl:
|
Open the file read-only so an inspection can never disturb a running crawl:
|
||||||
|
|||||||
@@ -0,0 +1,143 @@
|
|||||||
|
-- normalize-notes.sql
|
||||||
|
-- Replaces the repeated blog-name and type TEXT in Notes with integer IDs.
|
||||||
|
-- Reduces TL.db from ~207 MB to ~148 MB (-29%).
|
||||||
|
--
|
||||||
|
-- THIS IS A BREAKING SCHEMA CHANGE. There is no compatibility layer. Every
|
||||||
|
-- query in URLNotesGrabberCORE and Rolodex that names Notes.RootBlogName,
|
||||||
|
-- Notes.NoteBlogName or Notes.Type stops working the moment this runs, and
|
||||||
|
-- stays broken until those queries are rewritten. This was a deliberate choice
|
||||||
|
-- over a view-plus-triggers shim, which was measured to work but cost 194 ms ->
|
||||||
|
-- 321 ms on Rolodex's unfiltered Notes page.
|
||||||
|
--
|
||||||
|
-- TumblThree is unaffected. It touches only Blogs, and the column added to
|
||||||
|
-- Blogs here is additive.
|
||||||
|
--
|
||||||
|
-- HOW TO RUN (DB Browser for SQLite):
|
||||||
|
-- 1. Stop all three apps. Pause NextCloud sync.
|
||||||
|
-- 2. Back up TL.db.
|
||||||
|
-- 3. Execute SQL, paste this file, run. Then Write Changes.
|
||||||
|
-- 4. Tools > Compact Database (VACUUM). Nothing shrinks until this finishes.
|
||||||
|
|
||||||
|
--------------------------------------------------------------------------
|
||||||
|
-- The shape this produces
|
||||||
|
--------------------------------------------------------------------------
|
||||||
|
-- BlogNames(BlogId, BlogName) the ID authority: every name appearing in
|
||||||
|
-- Notes as either participant. 20,430 rows.
|
||||||
|
-- 12 of these have no Blogs row -- the
|
||||||
|
-- registry has never been a superset of the
|
||||||
|
-- engagement graph, and still is not.
|
||||||
|
--
|
||||||
|
-- NoteTypes(TypeId, Type) 5 rows. Fixed set, but written as a table
|
||||||
|
-- rather than a CHECK so a new type is an
|
||||||
|
-- INSERT and not a schema migration.
|
||||||
|
--
|
||||||
|
-- Notes(...Id columns...) integer FKs in place of text. WITHOUT ROWID,
|
||||||
|
-- same 5-column key in the same column order.
|
||||||
|
--
|
||||||
|
-- Blogs.BlogId NEW additive column. Lets Notes join Blogs in
|
||||||
|
-- one integer hop instead of going through
|
||||||
|
-- BlogNames and comparing text at the end.
|
||||||
|
-- NULL on the 168,202 blogs with no notes.
|
||||||
|
|
||||||
|
PRAGMA foreign_keys = off;
|
||||||
|
|
||||||
|
BEGIN;
|
||||||
|
|
||||||
|
--------------------------------------------------------------------------
|
||||||
|
-- STEP 1: the ID authority
|
||||||
|
--------------------------------------------------------------------------
|
||||||
|
CREATE TABLE BlogNames (
|
||||||
|
BlogId INTEGER PRIMARY KEY,
|
||||||
|
BlogName TEXT NOT NULL UNIQUE
|
||||||
|
);
|
||||||
|
|
||||||
|
INSERT INTO BlogNames (BlogName)
|
||||||
|
SELECT RootBlogName FROM Notes
|
||||||
|
UNION
|
||||||
|
SELECT NoteBlogName FROM Notes;
|
||||||
|
|
||||||
|
--------------------------------------------------------------------------
|
||||||
|
-- STEP 2: the type lookup
|
||||||
|
--------------------------------------------------------------------------
|
||||||
|
CREATE TABLE NoteTypes (
|
||||||
|
TypeId INTEGER PRIMARY KEY,
|
||||||
|
Type TEXT NOT NULL UNIQUE
|
||||||
|
);
|
||||||
|
|
||||||
|
-- IDs are assigned explicitly and must stay stable: they are stored in Notes.
|
||||||
|
INSERT INTO NoteTypes (TypeId, Type) VALUES
|
||||||
|
(1, 'like'),
|
||||||
|
(2, 'reblog'),
|
||||||
|
(3, 'reply'),
|
||||||
|
(4, 'posted'),
|
||||||
|
(5, 'post_attribution');
|
||||||
|
|
||||||
|
--------------------------------------------------------------------------
|
||||||
|
-- STEP 3: rebuild Notes with integer keys
|
||||||
|
--------------------------------------------------------------------------
|
||||||
|
-- Column order of the primary key is unchanged from the text version, so the
|
||||||
|
-- leading-column access patterns callers already rely on still hold:
|
||||||
|
-- (RootBlogId) and (RootBlogId, PostID) remain cheap prefixes.
|
||||||
|
CREATE TABLE NotesN (
|
||||||
|
RootBlogId INTEGER NOT NULL,
|
||||||
|
PostID INTEGER NOT NULL,
|
||||||
|
NoteBlogId INTEGER NOT NULL,
|
||||||
|
TimeStamp INTEGER NOT NULL,
|
||||||
|
TypeId INTEGER NOT NULL,
|
||||||
|
replyText TEXT,
|
||||||
|
DatetimeCrawled TEXT,
|
||||||
|
DateModified TEXT,
|
||||||
|
DateCreated TEXT,
|
||||||
|
IsActive INTEGER NOT NULL DEFAULT 1,
|
||||||
|
PRIMARY KEY (RootBlogId, PostID, TimeStamp, TypeId, NoteBlogId)
|
||||||
|
) WITHOUT ROWID;
|
||||||
|
|
||||||
|
-- Inner joins are safe here: BlogNames was just built from these very columns,
|
||||||
|
-- and NoteTypes covers all 5 values present. A row that failed to match would
|
||||||
|
-- be silently dropped, which is what the row-count check at the bottom is for.
|
||||||
|
INSERT INTO NotesN
|
||||||
|
SELECT r.BlogId, n.PostID, b.BlogId, n.TimeStamp, t.TypeId,
|
||||||
|
n.replyText, n.DatetimeCrawled, n.DateModified, n.DateCreated, n.IsActive
|
||||||
|
FROM Notes n
|
||||||
|
JOIN BlogNames r ON r.BlogName = n.RootBlogName
|
||||||
|
JOIN BlogNames b ON b.BlogName = n.NoteBlogName
|
||||||
|
JOIN NoteTypes t ON t.Type = n.Type;
|
||||||
|
|
||||||
|
DROP TABLE Notes;
|
||||||
|
ALTER TABLE NotesN RENAME TO Notes;
|
||||||
|
|
||||||
|
-- Replaces ix_NoteBlogName01. Renamed because it indexes a different column now.
|
||||||
|
CREATE INDEX ix_Notes_NoteBlogId ON Notes (NoteBlogId);
|
||||||
|
|
||||||
|
--------------------------------------------------------------------------
|
||||||
|
-- STEP 4: give Blogs the matching id
|
||||||
|
--------------------------------------------------------------------------
|
||||||
|
-- Additive: no existing column changes, so TumblThree's
|
||||||
|
-- "UPDATE Blogs SET IsActive = 0 ... WHERE BlogName = ?" is untouched.
|
||||||
|
ALTER TABLE Blogs ADD COLUMN BlogId INTEGER;
|
||||||
|
|
||||||
|
UPDATE Blogs
|
||||||
|
SET BlogId = (SELECT bn.BlogId FROM BlogNames bn WHERE bn.BlogName = Blogs.BlogName);
|
||||||
|
|
||||||
|
CREATE INDEX ix_Blogs_BlogId ON Blogs (BlogId);
|
||||||
|
|
||||||
|
COMMIT;
|
||||||
|
|
||||||
|
--------------------------------------------------------------------------
|
||||||
|
-- STEP 5: Write Changes, then Tools > Compact Database
|
||||||
|
--------------------------------------------------------------------------
|
||||||
|
-- From the CLI instead: sqlite3 TL.db "VACUUM;"
|
||||||
|
|
||||||
|
--------------------------------------------------------------------------
|
||||||
|
-- VERIFY
|
||||||
|
--------------------------------------------------------------------------
|
||||||
|
-- PRAGMA integrity_check; -- expect: ok
|
||||||
|
-- SELECT COUNT(*) FROM Notes; -- expect: 1182333, unchanged
|
||||||
|
-- SELECT COUNT(*) FROM BlogNames; -- expect: 20430
|
||||||
|
-- SELECT COUNT(*) FROM Blogs WHERE BlogId IS NOT NULL; -- expect: 20418
|
||||||
|
--
|
||||||
|
-- Losslessness was proven before this ran, by reconstructing the old text shape
|
||||||
|
-- from the new schema and diffing it against the original both ways:
|
||||||
|
-- SELECT COUNT(*) FROM (SELECT * FROM old.Notes EXCEPT SELECT * FROM Rebuilt);
|
||||||
|
-- SELECT COUNT(*) FROM (SELECT * FROM Rebuilt EXCEPT SELECT * FROM old.Notes);
|
||||||
|
-- Both returned 0 across all 1,182,333 rows and all 10 columns.
|
||||||
Reference in New Issue
Block a user