Compare commits
10
Commits
a28c5cc9ec
...
master
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
5af284a49f | ||
|
|
8fe2ffeb96 | ||
|
|
c8c43c4918 | ||
|
|
2948a4aff0 | ||
|
|
b8231d6a4c | ||
|
|
bbf05b3863 | ||
|
|
ea2afc9d35 | ||
|
|
7bf270e47d | ||
|
|
58d7b1d05e | ||
|
|
f41957fd2f |
@@ -31,7 +31,7 @@ dotnet run -- --test [blogname] [postID] # Test API for specific post
|
||||
- `--test [blogname] [postID]`: Test API note collection
|
||||
- `--posts`: Export post blogs to file
|
||||
- `--blogs`: Export blog list to file
|
||||
- `--collect [0|1] [datetime] [blogname]`: Collect notes for posts in DB. Optional `blogname` restricts the run to one blog (exact match), e.g. `--collect 1 zomb-eh`
|
||||
- `--collect [0|1] [datetime] [blogname]`: Collect notes for posts in DB. Optional `blogname` restricts the run to one blog (exact match), e.g. `--collect 1 zomb-eh`. Add `--force` to ignore the periodic re-collect cooldown so already-collected posts are re-queued immediately (mode 1 only). Add `--fromDate <datetime>` / `--toDate <datetime>` to only re-queue already-collected posts whose original PostDate is on/after / on/before that date (mode 1 only; either or both may be given; applies with or without `--force`)
|
||||
- `--blogsR`: Export reply blogs to file
|
||||
- `--blogsO [start] [stop]`: Export blogs within range
|
||||
|
||||
|
||||
@@ -49,28 +49,40 @@ say nothing about the item being fetched, so they must not be recorded as per-it
|
||||
|
||||
### `Notes` Stores Integer IDs, Not Names
|
||||
As of 2026-08-07 `Notes.RootBlogName`, `NoteBlogName` and `Type` are gone, replaced by
|
||||
`RootBlogId`, `NoteBlogId` and `TypeId` resolving through the `BlogNames` and `NoteTypes`
|
||||
lookup tables. There is no compatibility view — naming an old column is a hard SQLite
|
||||
error, so unlike `IsActive` this is a hard cut with no runtime probe. Full detail in
|
||||
`URLNotesGrabberCORE/TL.db.md`.
|
||||
`RootBlogId`, `NoteBlogId` and `TypeId`. Blog IDs resolve through `Blogs.BlogId`, and
|
||||
types through the `NoteTypes` lookup table. There is no compatibility view: naming an old
|
||||
column is a hard SQLite error, so unlike `IsActive` this is a hard cut with no runtime
|
||||
probe. Full detail in `URLNotesGrabberCORE/TL.db.md`.
|
||||
|
||||
- **Joining `Notes` to `Blogs` goes through `Blogs.BlogId`**, not `BlogNames`:
|
||||
`FROM Blogs B INNER JOIN Notes N ON N.NoteBlogId = B.BlogId`. Routing it through
|
||||
`BlogNames` adds a hop and ends in the text comparison the migration removed
|
||||
- **Joining `Notes` to `Posts` is the opposite** — `Posts` has only `BlogName`, so it must
|
||||
go through `BlogNames` (`GetRepliesWithFilledText`). This is the only such join
|
||||
- **Resolve a name by filtering the lookup, never by scanning `Notes`**:
|
||||
`WHERE NoteBlogId = (SELECT BlogId FROM BlogNames WHERE BlogName = @name)`. The subquery
|
||||
is a unique-index probe on 20k rows and does not show against the 1.18M-row table
|
||||
- **`AddNote` registers both blog names *and* the note type** with `INSERT OR IGNORE`
|
||||
before inserting, all in one transaction. `NoteTypes` is a table rather than a `CHECK`
|
||||
constraint precisely so an unseen type is an `INSERT`; without that registration it
|
||||
would resolve to `NULL` and fail the `NOT NULL` on `TypeId`, losing the note
|
||||
- **`Blogs.BlogId` is NULL on 168,202 of 188,620 rows** — every blog that has never
|
||||
appeared in a note. An inner join on it silently drops them. Correct for engagement
|
||||
queries, wrong for anything listing the registry
|
||||
- **IDs are stable and must never be renumbered.** They are stored in 1.18M `Notes` rows.
|
||||
A blog renamed upstream gets a new `BlogNames` row, not an edited one
|
||||
- **`Blogs.BlogId` is the only blog-ID authority (since 2026-09-28).** IDs used to live in a
|
||||
`BlogNames` table with an unmaintained copy in `Blogs.BlogId`. The copy drifted and hid
|
||||
12k blogs from `GetBlogs`, so `retire-blognames.sql` moved the authority into `Blogs`
|
||||
and **dropped `BlogNames` entirely**. There is no compatibility view, so naming it is
|
||||
`no such table`. Do not recreate it
|
||||
- **Joining `Notes` to `Blogs`**: `FROM Blogs B INNER JOIN Notes N ON N.NoteBlogId = B.BlogId`
|
||||
- **Joining `Notes` to `Posts` also goes through `Blogs`**, since `Posts` has only
|
||||
`BlogName`: `Posts P JOIN Blogs RB ON RB.BlogName = P.BlogName JOIN Notes N ON
|
||||
N.RootBlogId = RB.BlogId` (`GetRepliesWithFilledText`)
|
||||
- **Resolve a name by filtering `Blogs`, never by scanning `Notes`**:
|
||||
`WHERE NoteBlogId = (SELECT BlogId FROM Blogs WHERE BlogName = @name)`. The subquery is a
|
||||
primary-key probe and does not show against the 1.2M-row table
|
||||
- **`AddNote` registers both blogs *and* the note type** before inserting, all in one
|
||||
transaction. `RegisterBlog` does `INSERT OR IGNORE` into `Blogs`, then assigns
|
||||
`BlogId = MAX(BlogId) + 1` where it is NULL. Unlike `AddBlog`, it does not skip `deact`
|
||||
names, because a note by a deactivated blog still needs an ID. `NoteTypes` is a table
|
||||
rather than a `CHECK` constraint precisely so an unseen type is an `INSERT`. Without
|
||||
that registration a type would resolve to `NULL` and fail the `NOT NULL` on `TypeId`,
|
||||
losing the note
|
||||
- **Assigning a `BlogId` is bookkeeping and must not move `DateModified`**
|
||||
- **`Blogs.BlogId` is NULL on ~166k of ~199k rows**, every blog that has never appeared in
|
||||
a note. An inner join on it silently drops them. Correct for engagement queries, wrong
|
||||
for anything listing the registry. `ix_Blogs_BlogId` is `UNIQUE`, which allows many NULLs
|
||||
- **IDs are stable and must never be renumbered.** They are stored in 1.2M `Notes` rows.
|
||||
Triggers `trg_Blogs_BlogId_NoDelete` and `trg_Blogs_BlogId_Immutable` abort any
|
||||
`DELETE` of a `Blogs` row that has a `BlogId`, and any change to its `BlogId` or
|
||||
`BlogName`. A blog renamed upstream gets a new row. Remove a blog with `IsActive = 0`.
|
||||
These triggers are also what make `MAX(BlogId) + 1` safe: no ID can ever be freed for
|
||||
reuse
|
||||
- Prefer `TypeId = (SELECT TypeId FROM NoteTypes WHERE Type = 'reply')` over a hardcoded
|
||||
ID. A negated `TypeId NOT IN (SELECT …)` is only correct because `TypeId` is `NOT NULL`
|
||||
- Duplicate-key detection uses `IsNotesDuplicateKey`, which matches the constraint and the
|
||||
|
||||
+372
-66
@@ -1,59 +1,69 @@
|
||||
<?xml version="1.0" encoding="UTF-8"?><sqlb_project><db path="D:/NextCloud/C#/URLNotesGrabberCORE/URLNotesGrabberCORE/TL.db" readonly="0" foreign_keys="1" case_sensitive_like="0" temp_store="0" wal_autocheckpoint="1000" synchronous="2"/><attached/><window><main_tabs open="structure browser pragmas query" current="3"/></window><tab_structure><column_width id="0" width="300"/><column_width id="1" width="0"/><column_width id="2" width="100"/><column_width id="3" width="4486"/><column_width id="4" width="0"/><expanded_item id="0" parent="1"/><expanded_item id="1" parent="1"/><expanded_item id="2" parent="1"/><expanded_item id="3" parent="1"/></tab_structure><tab_browse><table title="Posts" custom_title="0" dock_id="4" table="4,5:mainPosts"/><dock_state state="000000ff00000000fd0000000100000002000005470000029afc0100000006fb000000160064006f0063006b00420072006f00770073006500310100000000000004a10000000000000000fb000000160064006f0063006b00420072006f00770073006500320100000000000004a10000000000000000fb000000160064006f0063006b00420072006f00770073006500330100000000000004a10000000000000000fb000000160064006f0063006b00420072006f00770073006500350100000000000005f40000000000000000fb000000160064006f0063006b00420072006f00770073006500340100000000000005470000011100fffffffb000000160064006f0063006b00420072006f00770073006500340100000000000005f40000000000000000000005470000000000000004000000040000000800000008fc00000000"/><default_encoding codec=""/><browse_table_settings><table schema="main" name="ApiKeyPoolMeta" show_row_id="0" encoding="" plot_x_axis="" unlock_view_pk="_rowid_" freeze_columns="0"><sort/><column_widths><column index="1" value="29"/><column index="2" value="64"/></column_widths><filter_values/><conditional_formats/><row_id_formats/><display_formats/><hidden_columns/><plot_y_axes/><global_filter/></table><table schema="main" name="Blogs" show_row_id="0" encoding="" plot_x_axis="" unlock_view_pk="_rowid_" freeze_columns="0"><sort><column index="7" mode="1"/></sort><column_widths><column index="1" value="257"/><column index="2" value="108"/><column index="3" value="63"/><column index="4" value="156"/><column index="5" value="60"/><column index="6" value="81"/><column index="7" value="85"/><column index="8" value="156"/><column index="9" value="156"/><column index="10" value="151"/><column index="11" value="125"/><column index="12" value="129"/></column_widths><filter_values><column index="4" value="1"/><column index="7" value=">2026-05-27 17:22:36"/></filter_values><conditional_formats/><row_id_formats/><display_formats/><hidden_columns/><plot_y_axes/><global_filter/></table><table schema="main" name="Notes" show_row_id="0" encoding="" plot_x_axis="" unlock_view_pk="_rowid_" freeze_columns="0"><sort/><column_widths><column index="1" value="198"/><column index="2" value="144"/><column index="3" value="251"/><column index="4" value="84"/><column index="5" value="53"/><column index="6" value="300"/><column index="7" value="116"/><column index="8" value="152"/><column index="9" value="89"/><column index="10" value="63"/></column_widths><filter_values/><conditional_formats/><row_id_formats/><display_formats/><hidden_columns/><plot_y_axes/><global_filter/></table><table schema="main" name="Posts" show_row_id="0" encoding="" plot_x_axis="" unlock_view_pk="_rowid_" freeze_columns="0"><sort><column index="14" mode="1"/></sort><column_widths><column index="1" value="236"/><column index="2" value="144"/><column index="3" value="126"/><column index="4" value="32"/><column index="5" value="32"/><column index="6" value="32"/><column index="7" value="32"/><column index="8" value="32"/><column index="9" value="0"/><column index="10" value="0"/><column index="11" value="0"/><column index="12" value="243"/><column index="13" value="300"/><column index="14" value="53"/><column index="15" value="37351"/><column index="16" value="300"/><column index="17" value="41"/><column index="18" value="75"/><column index="19" value="96"/><column index="20" value="300"/><column index="21" value="96"/><column index="22" value="300"/><column index="23" value="300"/><column index="24" value="548"/><column index="25" value="60"/><column index="26" value="213"/><column index="27" value="532"/><column index="28" value="152"/><column index="29" value="152"/><column index="30" value="69"/><column index="31" value="63"/></column_widths><filter_values><column index="2" value="0"/><column index="1" value="734568371821084672"/></filter_values><conditional_formats/><row_id_formats/><display_formats/><hidden_columns><column index="9" value="1"/><column index="10" value="1"/><column index="11" value="1"/></hidden_columns><plot_y_axes/><global_filter/></table></browse_table_settings></tab_browse><tab_sql><sql name="SQL 1">UPDATE Posts
|
||||
SET HasNotesGathered = 0
|
||||
WHERE (BlogName, PostID) IN (
|
||||
SELECT p.BlogName, p.PostID
|
||||
FROM Posts p
|
||||
WHERE p.HasNotesGathered = 1
|
||||
AND P.notesGatheredDatetime < 1774294520
|
||||
AND EXISTS (
|
||||
SELECT 1
|
||||
FROM Notes n
|
||||
WHERE n.PostID = p.PostID
|
||||
AND n.RootBlogId = (SELECT BlogId FROM BlogNames WHERE BlogName = p.BlogName)
|
||||
--AND n.TypeId NOT IN (SELECT TypeId FROM NoteTypes WHERE Type IN ('reblog', 'reply'))
|
||||
)
|
||||
ORDER BY P.PostDate ASC
|
||||
--LIMIT 500
|
||||
);</sql><sql name="Mark Blogs">select *
|
||||
<?xml version="1.0" encoding="UTF-8"?><sqlb_project><db path="C:/Users/jim/Nextcloud/C#/URLNotesGrabberCORE/URLNotesGrabberCORE/TL.db" readonly="0" foreign_keys="1" case_sensitive_like="0" temp_store="0" wal_autocheckpoint="1000" synchronous="2"/><attached/><window><main_tabs open="structure browser pragmas query" current="3"/></window><tab_structure><column_width id="0" width="300"/><column_width id="1" width="0"/><column_width id="2" width="100"/><column_width id="3" width="4486"/><column_width id="4" width="0"/><expanded_item id="0" parent="1"/><expanded_item id="1" parent="1"/><expanded_item id="2" parent="1"/><expanded_item id="3" parent="1"/></tab_structure><tab_browse><table title="Notes" custom_title="0" dock_id="4" table="4,5:mainNotes"/><dock_state state="000000ff00000000fd0000000100000002000005470000029efc0100000006fb000000160064006f0063006b00420072006f00770073006500310100000000000004a10000000000000000fb000000160064006f0063006b00420072006f00770073006500320100000000000004a10000000000000000fb000000160064006f0063006b00420072006f00770073006500330100000000000004a10000000000000000fb000000160064006f0063006b00420072006f00770073006500350100000000000005f40000000000000000fb000000160064006f0063006b00420072006f00770073006500340100000000000005470000011100fffffffb000000160064006f0063006b00420072006f00770073006500340100000000000005f40000000000000000000005470000000000000004000000040000000800000008fc00000000"/><default_encoding codec=""/><browse_table_settings><table schema="main" name="ApiKeyPoolMeta" show_row_id="0" encoding="" plot_x_axis="" unlock_view_pk="_rowid_" freeze_columns="0"><sort/><column_widths><column index="1" value="29"/><column index="2" value="64"/></column_widths><filter_values/><conditional_formats/><row_id_formats/><display_formats/><hidden_columns/><plot_y_axes/><global_filter/></table><table schema="main" name="Blogs" show_row_id="0" encoding="" plot_x_axis="" unlock_view_pk="_rowid_" freeze_columns="0"><sort><column index="7" mode="1"/></sort><column_widths><column index="1" value="257"/><column index="2" value="108"/><column index="3" value="63"/><column index="4" value="156"/><column index="5" value="60"/><column index="6" value="81"/><column index="7" value="85"/><column index="8" value="156"/><column index="9" value="156"/><column index="10" value="151"/><column index="11" value="125"/><column index="12" value="129"/></column_widths><filter_values><column index="4" value="1"/><column index="7" value=">2026-05-27 17:22:36"/></filter_values><conditional_formats/><row_id_formats/><display_formats/><hidden_columns/><plot_y_axes/><global_filter/></table><table schema="main" name="Notes" show_row_id="0" encoding="" plot_x_axis="" unlock_view_pk="_rowid_" freeze_columns="0"><sort><column index="4" mode="1"/></sort><column_widths><column index="1" value="81"/><column index="2" value="148"/><column index="3" value="83"/><column index="4" value="85"/><column index="5" value="56"/><column index="6" value="300"/><column index="7" value="156"/><column index="8" value="156"/><column index="9" value="156"/><column index="10" value="63"/></column_widths><filter_values><column index="2" value="4370"/></filter_values><conditional_formats/><row_id_formats/><display_formats/><hidden_columns/><plot_y_axes/><global_filter/></table><table schema="main" name="Posts" show_row_id="0" encoding="" plot_x_axis="" unlock_view_pk="_rowid_" freeze_columns="0"><sort><column index="14" mode="1"/></sort><column_widths><column index="1" value="241"/><column index="2" value="148"/><column index="3" value="126"/><column index="4" value="300"/><column index="5" value="75"/><column index="6" value="187"/><column index="7" value="159"/><column index="8" value="75"/><column index="9" value="0"/><column index="10" value="0"/><column index="11" value="0"/><column index="12" value="249"/><column index="13" value="300"/><column index="14" value="53"/><column index="15" value="300"/><column index="16" value="300"/><column index="17" value="41"/><column index="18" value="75"/><column index="19" value="96"/><column index="20" value="300"/><column index="21" value="96"/><column index="22" value="300"/><column index="23" value="300"/><column index="24" value="300"/><column index="25" value="60"/><column index="26" value="218"/><column index="27" value="300"/><column index="28" value="156"/><column index="29" value="156"/><column index="30" value="69"/><column index="31" value="63"/></column_widths><filter_values><column index="1" value="137735301451"/></filter_values><conditional_formats/><row_id_formats/><display_formats/><hidden_columns><column index="9" value="1"/><column index="10" value="1"/><column index="11" value="1"/></hidden_columns><plot_y_axes/><global_filter/></table></browse_table_settings></tab_browse><tab_sql><sql name="Mark Blogs">select *
|
||||
from Blogs
|
||||
--update blogs set HasBeenOutput = 1
|
||||
where HasBeenOutput = 0
|
||||
AND
|
||||
blogname in
|
||||
(
|
||||
'teaberrybee',
|
||||
'reddevilgoddesstoo',
|
||||
'waywardog13',
|
||||
'wzjustbrowsing-blog',
|
||||
'lewerta',
|
||||
'nudenymph',
|
||||
'caylachief'
|
||||
|
||||
('udontn33dh1m',
|
||||
'tyrantsxblood',
|
||||
'sentry-34',
|
||||
'deathcabforfrankie',
|
||||
'abheith-sasta',
|
||||
'kuwaiikittenghost',
|
||||
'kansasmud',
|
||||
'03diesel',
|
||||
'itzameallieee',
|
||||
'fireball-temptations',
|
||||
'mamaisamess',
|
||||
'906raised-and-dogobsessed',
|
||||
'the-queerist-wolf',
|
||||
'counting-corpsess',
|
||||
'aqueenbby',
|
||||
'maybememoriesx',
|
||||
'queenofnevers',
|
||||
'obsidian-psyche',
|
||||
'lilmissellexo',
|
||||
'alittlebunny95',
|
||||
'rage--and--grace',
|
||||
'savage-deniz',
|
||||
'daddyspuddleprincess',
|
||||
'littledefenstration',
|
||||
'bearded-snorlax',
|
||||
'thosesummerskiess',
|
||||
'tubadtoph',
|
||||
'lieutenant-dan-ice-cream',
|
||||
'brittvnybitch',
|
||||
'a-smol-gayologist',
|
||||
'sum1random',
|
||||
'samsternelly',
|
||||
'littlemouseylauren',
|
||||
'princessleiaorgasma',
|
||||
'bloodstaineddkisses',
|
||||
'letsfacerealitybabe',
|
||||
'x--marks--thespot',
|
||||
'space-and-suffering',
|
||||
'rinarootski',
|
||||
'thiccandtired',
|
||||
'fvcking-scvmbag',
|
||||
'fullblownwizard',
|
||||
'bigjewface',
|
||||
'unleash-the-krayken',
|
||||
'bumpintheroad',
|
||||
'liltexasjedii',
|
||||
'nawtydude',
|
||||
'queenpeachqueen',
|
||||
'the-clansman',
|
||||
'balmain-bxtch'
|
||||
)</sql><sql name="New Notes">select P.slug, N.replyText, rbn.BlogName as RootBlogName, n.PostID, nbn.BlogName || '.tumblr.com' as NoteBlogName, DatetimeCrawled, TimeStamp, nt.Type, rbn.BlogName || '.tumblr.com/post/' || n.postid, datetime(timestamp, 'unixepoch')
|
||||
from Notes N
|
||||
inner join Posts P on p.PostID = n.PostID
|
||||
inner join BlogNames rbn on rbn.BlogId = n.RootBlogId
|
||||
inner join BlogNames nbn on nbn.BlogId = n.NoteBlogId
|
||||
inner join Blogs rbn on rbn.BlogId = n.RootBlogId
|
||||
inner join Blogs nbn on nbn.BlogId = n.NoteBlogId
|
||||
inner join NoteTypes nt on nt.TypeId = n.TypeId
|
||||
where
|
||||
DatetimeCrawled > '2026-08-07 11:47:22' and nt.Type like 'r%'
|
||||
and P.IsActive = 1
|
||||
order by n.DatetimeCrawled</sql><sql name="Pull Blogs">SELECT distinct
|
||||
'''' || blogname || ''',',
|
||||
blogs.*
|
||||
, blogname || '.tumblr.com'
|
||||
FROM
|
||||
Blogs
|
||||
inner JOIN
|
||||
Notes on notes.noteBlogId = blogs.BlogId
|
||||
inner JOIN
|
||||
NoteTypes on NoteTypes.TypeId = Notes.TypeId
|
||||
WHERE
|
||||
HasBeenOutput = 0 and NoteTypes.Type = 'reblog'
|
||||
order by
|
||||
NoteTypes.Type desc,
|
||||
DateAdded desc
|
||||
LIMIT 100;</sql><sql name="SQL 7">WITH ReplyCounts AS (
|
||||
order by n.DatetimeCrawled</sql><sql name="SQL 7">WITH ReplyCounts AS (
|
||||
SELECT
|
||||
NoteBlogId,
|
||||
COUNT(DISTINCT replyText) AS DistinctReplyCount
|
||||
@@ -68,13 +78,13 @@ SELECT
|
||||
c.DistinctReplyCount
|
||||
FROM Notes n
|
||||
JOIN ReplyCounts c ON n.NoteBlogId = c.NoteBlogId
|
||||
JOIN BlogNames rbn ON rbn.BlogId = n.RootBlogId
|
||||
JOIN BlogNames nbn ON nbn.BlogId = n.NoteBlogId
|
||||
JOIN Blogs rbn ON rbn.BlogId = n.RootBlogId
|
||||
JOIN Blogs nbn ON nbn.BlogId = n.NoteBlogId
|
||||
JOIN NoteTypes t ON t.TypeId = n.TypeId
|
||||
where replyText <> '.' and t.Type <> 'reply'
|
||||
--AND nbn.BlogName NOT IN ( 'roadblocker21', 'thesaddemon666', 'edwardabbeyhoffman', 'tattedsoldier20', 'zomb-eh', 'animalistic13', 'indken', 'maccloud1592',
|
||||
--'moss-wizard', 'supertrucker12682', 'exploringthrupics', 'padeyepete' )
|
||||
order by c.DistinctReplyCount desc, nbn.BlogName, n.DateModified desc, replyText, rbn.BlogName, PostID</sql><sql name="Collect">WITH PostsWithCount AS ( SELECT P.BlogName, P.PostID, 1925013599 AS LatestNoteTimestamp, P.NotesGatheredDateTime, COUNT(P.PostID) OVER(PARTITION BY P.BlogName) AS CNT, P.HasNotesGathered, P.NotFound, P.PostDate FROM Posts P WHERE COALESCE(P.IsActive, 1) = 1 ), Unioned AS ( SELECT BlogName, PostID, LatestNoteTimestamp, NotesGatheredDateTime, CNT, PostDate FROM PostsWithCount WHERE NotFound = 0 AND HasNotesGathered = 0 UNION SELECT BlogName, PostID, LatestNoteTimestamp, NotesGatheredDateTime, CNT, PostDate FROM PostsWithCount WHERE BlogName = 'zomb-eh' AND NotFound = 0 AND NotesGatheredDateTime < unixepoch('now', 'localtime', '-3 days') ) SELECT U.BlogName, U.PostID, U.LatestNoteTimestamp, U.NotesGatheredDateTime, U.CNT FROM Unioned U WHERE (U.NotesGatheredDateTime < 1786134037 OR U.NotesGatheredDateTime IS NULL) ORDER BY U.NotesGatheredDateTime, U.PostDate DESC, U.BlogName, U.PostID;</sql><sql name="Del Posts">delete from posts where postid in
|
||||
order by c.DistinctReplyCount desc, nbn.BlogName, n.DateModified desc, replyText, rbn.BlogName, PostID</sql><sql name="Collect">WITH PostsWithCount AS ( SELECT P.BlogName, P.PostID, 1925013599 AS LatestNoteTimestamp, P.NotesGatheredDateTime, COUNT(P.PostID) OVER(PARTITION BY P.BlogName) AS CNT, P.HasNotesGathered, P.NotFound, P.PostDate FROM Posts P WHERE COALESCE(P.IsActive, 1) = 1 ), Unioned AS ( SELECT BlogName, PostID, LatestNoteTimestamp, NotesGatheredDateTime, CNT, PostDate FROM PostsWithCount WHERE NotFound = 0 AND HasNotesGathered = 0 UNION SELECT BlogName, PostID, LatestNoteTimestamp, NotesGatheredDateTime, CNT, PostDate FROM PostsWithCount WHERE BlogName = 'zomb-eh' AND NotFound = 0 AND NotesGatheredDateTime < unixepoch('now', 'localtime', '-3 days') ) SELECT U.BlogName, U.PostID, U.LatestNoteTimestamp, U.NotesGatheredDateTime, U.CNT FROM Unioned U WHERE (U.NotesGatheredDateTime < 1787237598 OR U.NotesGatheredDateTime IS NULL) ORDER BY U.NotesGatheredDateTime, U.PostDate DESC, U.BlogName, U.PostID;</sql><sql name="Del Posts">delete from posts where postid in
|
||||
(
|
||||
'741662499571728384',
|
||||
178892849664,
|
||||
@@ -82,28 +92,324 @@ order by c.DistinctReplyCount desc, nbn.BlogName, n.DateModified desc, replyText
|
||||
177012868749,
|
||||
169950081964,
|
||||
755440787056099328
|
||||
)</sql><sql name="notes NO post*">select *
|
||||
)</sql><sql name="notes NO post">select *
|
||||
-- delete
|
||||
from notes
|
||||
where postid not in (select distinct postid from posts where IsActive = 1)</sql><sql name="SQL 9">SELECT
|
||||
*
|
||||
FROM
|
||||
POSTS P
|
||||
WHERE
|
||||
P.ByLikes = 1
|
||||
AND
|
||||
P.DateCreated > '2026-05-26 17:47:32'
|
||||
ORDER BY
|
||||
P.DateCreated desc</sql><sql name="SQL 13">update posts set IsActive = 0 where blogname IN ( 'shoebiedoo', 'redheaded-girlygirl', 'xlittle-ghost' )</sql><sql name="SQL 14*">update Posts
|
||||
set IsActive = 0
|
||||
where postid in
|
||||
(
|
||||
|
||||
|
||||
'731937314675310592'
|
||||
|
||||
where postid not in (select distinct postid from posts where IsActive = 1)</sql><sql name="Pull Blogs">-- ============================================================================
|
||||
-- blogs-added-after-august-2026-with-reblog-or-reply.sql
|
||||
--
|
||||
-- Purpose: Of all blogs in Blogs, find the ones that show up in Notes as the
|
||||
-- engager (NoteBlogId) on a 'reblog' or 'reply' note, sorted by
|
||||
-- DateAdded. (Originally scoped to "added after August 2026" --
|
||||
-- that cutoff is now removed per request; QUERY 2 shows how to put
|
||||
-- a date floor back if needed.)
|
||||
--
|
||||
-- Read-only. No INSERT/UPDATE/DELETE/DDL anywhere in this file.
|
||||
--
|
||||
-- How to use (DB Browser for SQLite):
|
||||
-- 1. File > Open Database -> TL.db
|
||||
-- 2. Execute SQL tab. Each QUERY below is independent; Ctrl+Enter runs just
|
||||
-- the one your cursor is in.
|
||||
--
|
||||
-- The join, once:
|
||||
-- "Added after August 2026" filters Blogs.DateAdded. "Has a reblog/reply
|
||||
-- note" means the blog is the engager, which is NoteBlogId -- not
|
||||
-- RootBlogId, which is the blog that *owns* the post being reacted to
|
||||
-- (see find-notes-on-inactive-posts.sql for that side). Per TL.db.md, the
|
||||
-- Blogs<->Notes join is a single integer hop and should not be routed
|
||||
-- through Blogs:
|
||||
-- Blogs.BlogId = Notes.NoteBlogId
|
||||
-- EXISTS is used rather than a JOIN so a blog with many qualifying notes
|
||||
-- still contributes one output row.
|
||||
--
|
||||
-- Excluding notes on an inactive post: same shape as
|
||||
-- find-notes-on-inactive-posts.sql -- Notes only carries RootBlogId (an
|
||||
-- integer), so reaching Posts.IsActive needs the one text hop the rest of
|
||||
-- this file avoids: RootBlogId -> Blogs.BlogName = Posts.BlogName,
|
||||
-- matched on PostID. It's a LEFT JOIN, not an inner one: only 3,867 of
|
||||
-- 20,430 root/engager blogs have any stored Posts rows at all (TL.db.md),
|
||||
-- so most reblog/reply notes have no Posts row to check and must be kept,
|
||||
-- not dropped by an inner join. COALESCE(p.IsActive, 1) = 1 keeps a note
|
||||
-- unless its post is explicitly IsActive = 0 -- NULL (no Posts row, or a
|
||||
-- stored row with no flag written) means live, per the schema's own
|
||||
-- convention (TL.db.md, "Posts.IsActive and Notes.IsActive"). This is a
|
||||
-- big filter in practice: of the blogs that qualified before it, most
|
||||
-- have every one of their reblog/reply notes pointing at a since-removed
|
||||
-- post, not just some -- verified against the live data, not assumed.
|
||||
--
|
||||
-- On DateAdded: this column is not written consistently -- most rows hold
|
||||
-- ISO 'yyyy-MM-dd HH:mm:ss', but 17k+ hold US 'M/d/yy' from a 2025 bulk
|
||||
-- import (see TL.db.md, "DateAdded is not written consistently"). As text,
|
||||
-- those two shapes do not sort or compare against each other correctly, so
|
||||
-- QUERY 0 normalises both to an ISO date before filtering. In the live data
|
||||
-- every US-format row predates August 2026 anyway (only '12/23/25' and
|
||||
-- '12/24/25' occur), so this makes no difference to the current answer --
|
||||
-- it's here so the query stays correct if that ever changes.
|
||||
-- ============================================================================
|
||||
|
||||
|
||||
-- ----------------------------------------------------------------------------
|
||||
-- QUERY 0 / THE ANSWER -- one row per qualifying blog, sorted by DateAdded
|
||||
-- descending (normalised -- see the note above). No date cutoff, but now
|
||||
-- scoped to HasBeenOutput = 0 AND IsActive = 1. 4,739 rows in the live
|
||||
-- data.
|
||||
--
|
||||
-- earliest_reblog_or_reply_utc is the MIN(TimeStamp) among this blog's
|
||||
-- reblog-or-reply notes (either type counts -- see the column name).
|
||||
-- Getting this meant switching QUERY 0 from EXISTS to an inner JOIN +
|
||||
-- GROUP BY: EXISTS can only tell you a qualifying row is present, not
|
||||
-- aggregate over which ones. No CASE is needed inside the MIN() because
|
||||
-- the WHERE below already restricts the joined rows to reblog/reply, so
|
||||
-- every row a blog brings into the aggregate is one this column should
|
||||
-- consider. A blog appears exactly once, same as before, and this column
|
||||
-- is never NULL for a row that's in the result at all (an earlier
|
||||
-- revision aggregated reblog only, which left it NULL for the 181 blogs
|
||||
-- that had replies but no reblogs).
|
||||
-- ----------------------------------------------------------------------------
|
||||
WITH BlogsSplit AS (
|
||||
SELECT
|
||||
b.BlogId,
|
||||
b.BlogName,
|
||||
b.DateAdded,
|
||||
CASE WHEN b.DateAdded LIKE '____-__-__%' THEN 1 ELSE 0 END AS IsIso,
|
||||
-- for the US 'M/d/yy' shape only: everything after the first '/'
|
||||
substr(b.DateAdded, instr(b.DateAdded, '/') + 1) AS RestAfterMonth
|
||||
FROM Blogs b
|
||||
WHERE b.BlogId IS NOT NULL and HasBeenOutput = 0 and IsActive = 1 -- a blog can only match Notes if it has one
|
||||
),
|
||||
BlogsNorm AS (
|
||||
SELECT
|
||||
BlogId,
|
||||
BlogName,
|
||||
DateAdded,
|
||||
CASE
|
||||
WHEN IsIso = 1 THEN date(DateAdded)
|
||||
ELSE date(
|
||||
'20' || substr(RestAfterMonth, instr(RestAfterMonth, '/') + 1) || '-' ||
|
||||
substr('00' || substr(DateAdded, 1, instr(DateAdded, '/') - 1), -2) || '-' ||
|
||||
substr('00' || substr(RestAfterMonth, 1, instr(RestAfterMonth, '/') - 1), -2)
|
||||
)
|
||||
END AS DateAddedNorm
|
||||
FROM BlogsSplit
|
||||
)
|
||||
SELECT
|
||||
bn.BlogId,
|
||||
bn.BlogName,
|
||||
bn.DateAdded,
|
||||
bn.DateAddedNorm,
|
||||
datetime(MIN(n.TimeStamp), 'unixepoch') AS earliest_reblog_or_reply_utc
|
||||
FROM BlogsNorm bn
|
||||
JOIN Notes n ON n.NoteBlogId = bn.BlogId
|
||||
JOIN NoteTypes t ON t.TypeId = n.TypeId
|
||||
JOIN Blogs root_bn ON root_bn.BlogId = n.RootBlogId
|
||||
LEFT JOIN Posts p ON p.BlogName = root_bn.BlogName
|
||||
AND p.PostID = n.PostID
|
||||
WHERE t.Type IN ('reblog')--, 'reply')
|
||||
AND COALESCE(p.IsActive, 1) = 1 -- exclude notes on a post explicitly marked removed
|
||||
GROUP BY bn.BlogId, bn.BlogName, bn.DateAdded, bn.DateAddedNorm
|
||||
ORDER BY bn.DateAddedNorm desc;
|
||||
|
||||
</sql><current_tab id="7"/></tab_sql></sqlb_project>
|
||||
|
||||
-- ----------------------------------------------------------------------------
|
||||
-- QUERY 1 -- same answer, with a per-blog breakdown of which type(s) fired
|
||||
-- and how many. Useful once QUERY 0 has rows; redundant while it's empty.
|
||||
-- ----------------------------------------------------------------------------
|
||||
-- WITH BlogsSplit AS ( ... ), BlogsNorm AS ( ... ) -- reuse the CTEs above
|
||||
--
|
||||
-- SELECT
|
||||
-- bn.BlogId,
|
||||
-- bn.BlogName,
|
||||
-- bn.DateAddedNorm,
|
||||
-- SUM(CASE WHEN t.Type = 'reblog' THEN 1 ELSE 0 END) AS reblog_count,
|
||||
-- SUM(CASE WHEN t.Type = 'reply' THEN 1 ELSE 0 END) AS reply_count
|
||||
-- FROM BlogsNorm bn
|
||||
-- JOIN Notes n ON n.NoteBlogId = bn.BlogId
|
||||
-- JOIN NoteTypes t ON t.TypeId = n.TypeId
|
||||
-- WHERE t.Type IN ('reblog', 'reply')
|
||||
-- GROUP BY bn.BlogId, bn.BlogName, bn.DateAddedNorm
|
||||
-- ORDER BY bn.DateAddedNorm;
|
||||
|
||||
|
||||
-- ----------------------------------------------------------------------------
|
||||
-- QUERY 2 -- put a date floor back, if wanted later.
|
||||
-- Same as QUERY 0, with one extra line in the outer WHERE:
|
||||
-- AND bn.DateAddedNorm >= '2026-09-01' -- or whatever cutoff
|
||||
-- ----------------------------------------------------------------------------
|
||||
-- ============================================================================
|
||||
-- blogs-added-after-august-2026-with-reblog-or-reply.sql
|
||||
--
|
||||
-- Purpose: Of all blogs in Blogs, find the ones that show up in Notes as the
|
||||
-- engager (NoteBlogId) on a 'reblog' or 'reply' note, sorted by
|
||||
-- DateAdded. (Originally scoped to "added after August 2026" --
|
||||
-- that cutoff is now removed per request; QUERY 2 shows how to put
|
||||
-- a date floor back if needed.)
|
||||
--
|
||||
-- Read-only. No INSERT/UPDATE/DELETE/DDL anywhere in this file.
|
||||
--
|
||||
-- How to use (DB Browser for SQLite):
|
||||
-- 1. File > Open Database -> TL.db
|
||||
-- 2. Execute SQL tab. Each QUERY below is independent; Ctrl+Enter runs just
|
||||
-- the one your cursor is in.
|
||||
--
|
||||
-- The join, once:
|
||||
-- "Added after August 2026" filters Blogs.DateAdded. "Has a reblog/reply
|
||||
-- note" means the blog is the engager, which is NoteBlogId -- not
|
||||
-- RootBlogId, which is the blog that *owns* the post being reacted to
|
||||
-- (see find-notes-on-inactive-posts.sql for that side). Per TL.db.md, the
|
||||
-- Blogs<->Notes join is a single integer hop and should not be routed
|
||||
-- through Blogs:
|
||||
-- Blogs.BlogId = Notes.NoteBlogId
|
||||
-- EXISTS is used rather than a JOIN so a blog with many qualifying notes
|
||||
-- still contributes one output row.
|
||||
--
|
||||
-- Excluding notes on an inactive post: same shape as
|
||||
-- find-notes-on-inactive-posts.sql -- Notes only carries RootBlogId (an
|
||||
-- integer), so reaching Posts.IsActive needs the one text hop the rest of
|
||||
-- this file avoids: RootBlogId -> Blogs.BlogName = Posts.BlogName,
|
||||
-- matched on PostID. It's a LEFT JOIN, not an inner one: only 3,867 of
|
||||
-- 20,430 root/engager blogs have any stored Posts rows at all (TL.db.md),
|
||||
-- so most reblog/reply notes have no Posts row to check and must be kept,
|
||||
-- not dropped by an inner join. COALESCE(p.IsActive, 1) = 1 keeps a note
|
||||
-- unless its post is explicitly IsActive = 0 -- NULL (no Posts row, or a
|
||||
-- stored row with no flag written) means live, per the schema's own
|
||||
-- convention (TL.db.md, "Posts.IsActive and Notes.IsActive"). This is a
|
||||
-- big filter in practice: of the blogs that qualified before it, most
|
||||
-- have every one of their reblog/reply notes pointing at a since-removed
|
||||
-- post, not just some -- verified against the live data, not assumed.
|
||||
--
|
||||
-- On DateAdded: this column is not written consistently -- most rows hold
|
||||
-- ISO 'yyyy-MM-dd HH:mm:ss', but 17k+ hold US 'M/d/yy' from a 2025 bulk
|
||||
-- import (see TL.db.md, "DateAdded is not written consistently"). As text,
|
||||
-- those two shapes do not sort or compare against each other correctly, so
|
||||
-- QUERY 0 normalises both to an ISO date before filtering. In the live data
|
||||
-- every US-format row predates August 2026 anyway (only '12/23/25' and
|
||||
-- '12/24/25' occur), so this makes no difference to the current answer --
|
||||
-- it's here so the query stays correct if that ever changes.
|
||||
-- ============================================================================
|
||||
|
||||
|
||||
-- ----------------------------------------------------------------------------
|
||||
-- QUERY 0 / THE ANSWER -- one row per qualifying blog, sorted by
|
||||
-- earliest_reblog_or_reply_utc then DateAdded descending (normalised --
|
||||
-- see the note above). No date cutoff, but scoped to HasBeenOutput = 0
|
||||
-- AND IsActive = 1, and now excluding notes on a removed post (see the
|
||||
-- header note above). 1,804 rows in the live data as of this revision --
|
||||
-- down from 4,396 just before this exclusion was added, because most of
|
||||
-- the blogs that dropped out had *every* reblog/reply note pointing at a
|
||||
-- now-inactive post, not just some (the number moves between runs
|
||||
-- regardless -- crawling and output flip HasBeenOutput/IsActive on live
|
||||
-- rows).
|
||||
--
|
||||
-- earliest_reblog_or_reply_utc is the earliest TimeStamp among this
|
||||
-- blog's reblog-or-reply notes (either type counts -- see the column
|
||||
-- name); earliest_reblog_or_reply_postid and _root_blogid identify that
|
||||
-- specific note's post: PostID + RootBlogId together, not PostID alone --
|
||||
-- see TL.db.md ("345 post IDs exist under more than one blog"), same
|
||||
-- caution as in find-notes-on-inactive-posts.sql. Resolve RootBlogId to a
|
||||
-- name via Blogs (or Blogs, tolerating a miss) if you need it.
|
||||
--
|
||||
-- Getting "which note" rather than just "when" doesn't fit a plain
|
||||
-- MIN()/GROUP BY -- an aggregate can tell you the earliest value but not
|
||||
-- which row it came from. EarliestNote instead ranks each blog's
|
||||
-- reblog/reply notes with ROW_NUMBER() OVER (PARTITION BY NoteBlogId
|
||||
-- ORDER BY TimeStamp), and QUERY 0 takes rn = 1. The ORDER BY carries a
|
||||
-- PostID tiebreak because (NoteBlogId, TimeStamp) is not unique in this
|
||||
-- data -- ties exist (e.g. NoteBlogId 12 has 7 notes at the same
|
||||
-- TimeStamp) -- so without a tiebreak the "earliest" pick would be
|
||||
-- arbitrary among ties rather than deterministic.
|
||||
--
|
||||
-- EarliestNote also excludes notes on an inactive post before ranking
|
||||
-- (see the header note above), so "earliest" means earliest surviving
|
||||
-- note, not earliest overall -- a blog whose true-earliest note pointed
|
||||
-- at a since-removed post now surfaces its next-earliest live one
|
||||
-- instead. Applying the exclusion here, not as a filter on QUERY 0's
|
||||
-- final rows, matters: filtering after ROW_NUMBER would have picked the
|
||||
-- removed-post note as rn = 1 and then dropped the whole row instead of
|
||||
-- promoting the next candidate.
|
||||
-- ----------------------------------------------------------------------------
|
||||
WITH BlogsSplit AS (
|
||||
SELECT
|
||||
b.BlogId,
|
||||
b.BlogName,
|
||||
b.DateAdded,
|
||||
CASE WHEN b.DateAdded LIKE '____-__-__%' THEN 1 ELSE 0 END AS IsIso,
|
||||
-- for the US 'M/d/yy' shape only: everything after the first '/'
|
||||
substr(b.DateAdded, instr(b.DateAdded, '/') + 1) AS RestAfterMonth
|
||||
FROM Blogs b
|
||||
WHERE b.BlogId IS NOT NULL and HasBeenOutput = 0 and IsActive = 1 -- a blog can only match Notes if it has one
|
||||
),
|
||||
BlogsNorm AS (
|
||||
SELECT
|
||||
BlogId,
|
||||
BlogName,
|
||||
DateAdded,
|
||||
CASE
|
||||
WHEN IsIso = 1 THEN date(DateAdded)
|
||||
ELSE date(
|
||||
'20' || substr(RestAfterMonth, instr(RestAfterMonth, '/') + 1) || '-' ||
|
||||
substr('00' || substr(DateAdded, 1, instr(DateAdded, '/') - 1), -2) || '-' ||
|
||||
substr('00' || substr(RestAfterMonth, 1, instr(RestAfterMonth, '/') - 1), -2)
|
||||
)
|
||||
END AS DateAddedNorm
|
||||
FROM BlogsSplit
|
||||
),
|
||||
EarliestNote AS (
|
||||
SELECT
|
||||
n.NoteBlogId,
|
||||
n.RootBlogId,
|
||||
n.PostID,
|
||||
n.TimeStamp,
|
||||
ROW_NUMBER() OVER (
|
||||
PARTITION BY n.NoteBlogId
|
||||
ORDER BY n.TimeStamp ASC, n.PostID ASC
|
||||
) AS rn
|
||||
FROM Notes n
|
||||
JOIN NoteTypes t ON t.TypeId = n.TypeId
|
||||
JOIN Blogs root_bn ON root_bn.BlogId = n.RootBlogId
|
||||
LEFT JOIN Posts p ON p.BlogName = root_bn.BlogName
|
||||
AND p.PostID = n.PostID
|
||||
WHERE t.Type IN ('reblog')--, 'reply')
|
||||
AND n.NoteBlogId IN (SELECT BlogId FROM BlogsNorm) -- scope the window to blogs we care about
|
||||
AND COALESCE(p.IsActive, 1) = 1 -- exclude notes on a post explicitly marked removed
|
||||
)
|
||||
SELECT
|
||||
bn.BlogId,
|
||||
bn.BlogName || '.tumblr.com',
|
||||
'''' || bn.blogname || ''',',
|
||||
bn.DateAdded,
|
||||
bn.DateAddedNorm,
|
||||
datetime(en.TimeStamp, 'unixepoch') AS earliest_reblog_or_reply_utc,
|
||||
en.PostID AS earliest_reblog_or_reply_postid,
|
||||
en.RootBlogId AS earliest_reblog_or_reply_root_blogid
|
||||
FROM BlogsNorm bn
|
||||
JOIN EarliestNote en ON en.NoteBlogId = bn.BlogId AND en.rn = 1
|
||||
ORDER BY earliest_reblog_or_reply_utc, bn.DateAddedNorm desc
|
||||
limit 50;
|
||||
|
||||
|
||||
-- ----------------------------------------------------------------------------
|
||||
-- QUERY 1 -- same answer, with a per-blog breakdown of which type(s) fired
|
||||
-- and how many. Useful once QUERY 0 has rows; redundant while it's empty.
|
||||
-- ----------------------------------------------------------------------------
|
||||
-- WITH BlogsSplit AS ( ... ), BlogsNorm AS ( ... ) -- reuse the CTEs above
|
||||
--
|
||||
-- SELECT
|
||||
-- bn.BlogId,
|
||||
-- bn.BlogName,
|
||||
-- bn.DateAddedNorm,
|
||||
-- SUM(CASE WHEN t.Type = 'reblog' THEN 1 ELSE 0 END) AS reblog_count,
|
||||
-- SUM(CASE WHEN t.Type = 'reply' THEN 1 ELSE 0 END) AS reply_count
|
||||
-- FROM BlogsNorm bn
|
||||
-- JOIN Notes n ON n.NoteBlogId = bn.BlogId
|
||||
-- JOIN NoteTypes t ON t.TypeId = n.TypeId
|
||||
-- WHERE t.Type IN ('reblog', 'reply')
|
||||
-- GROUP BY bn.BlogId, bn.BlogName, bn.DateAddedNorm
|
||||
-- ORDER BY bn.DateAddedNorm;
|
||||
|
||||
|
||||
-- ----------------------------------------------------------------------------
|
||||
-- QUERY 2 -- put a date floor back, if wanted later.
|
||||
-- Same as QUERY 0, with one extra line in the outer WHERE:
|
||||
-- AND bn.DateAddedNorm >= '2026-09-01' -- or whatever cutoff
|
||||
-- ----------------------------------------------------------------------------
|
||||
</sql><current_tab id="0"/></tab_sql></sqlb_project>
|
||||
|
||||
@@ -15,9 +15,9 @@ ORDER BY
|
||||
FROM
|
||||
Notes N
|
||||
inner JOIN
|
||||
BlogNames rbn on rbn.BlogId = N.RootBlogId
|
||||
Blogs rbn on rbn.BlogId = N.RootBlogId
|
||||
inner JOIN
|
||||
BlogNames nbn on nbn.BlogId = N.NoteBlogId
|
||||
Blogs nbn on nbn.BlogId = N.NoteBlogId
|
||||
inner JOIN
|
||||
NoteTypes t on t.TypeId = N.TypeId
|
||||
inner JOIN
|
||||
|
||||
@@ -216,18 +216,19 @@ namespace URLNotesGrabberCORE
|
||||
#region Notes integer schema
|
||||
|
||||
// Notes stopped storing names on 2026-08-07: RootBlogName/NoteBlogName/Type became
|
||||
// RootBlogId/NoteBlogId/TypeId, resolved through BlogNames and NoteTypes. There is no
|
||||
// RootBlogId/NoteBlogId/TypeId, resolved through Blogs.BlogId and NoteTypes. There is no
|
||||
// compatibility view -- a query naming an old column fails outright, so this is a hard
|
||||
// cut rather than an optional column like IsActive. See TL.db.md.
|
||||
//
|
||||
// Blogs.BlogId is the only ID authority as of 2026-09-28; the BlogNames table that used
|
||||
// to hold the IDs is gone. Every name in Notes has a Blogs row, created by RegisterBlog.
|
||||
//
|
||||
// Two shapes recur below and are spelled out inline rather than hidden behind a helper,
|
||||
// so that every statement reads as the SQL it actually runs:
|
||||
// (SELECT BlogId FROM BlogNames WHERE BlogName = @name) -- unique-index probe, 20k rows
|
||||
// (SELECT BlogId FROM Blogs WHERE BlogName = @name) -- primary-key probe
|
||||
// (SELECT TypeId FROM NoteTypes WHERE Type = 'reply') -- 5 rows, effectively free
|
||||
// Joining Notes to Blogs is the one case that must NOT route through BlogNames: Blogs
|
||||
// carries its own BlogId, so N.NoteBlogId = B.BlogId is a single integer hop. Joining
|
||||
// Notes to Posts is the opposite case -- Posts has only BlogName, so it has to go
|
||||
// through BlogNames.
|
||||
// Joining Notes to Blogs is N.NoteBlogId = B.BlogId, a single hop on the unique index.
|
||||
// Joining Notes to Posts goes through Blogs too -- Posts has only BlogName.
|
||||
|
||||
/// <summary>
|
||||
/// True when the exception is a duplicate-key collision on Notes. The message embeds the
|
||||
@@ -242,14 +243,30 @@ namespace URLNotesGrabberCORE
|
||||
}
|
||||
|
||||
/// <summary>
|
||||
/// Gives a blog name an ID if it does not have one. No read-back and no round trip -- a
|
||||
/// name that is already registered keeps the ID that 1.18M Notes rows point at.
|
||||
/// Ensures a blog has a Blogs row and a BlogId, so a note can point at it. No read-back
|
||||
/// and no round trip -- a blog that already has an ID keeps the one Notes rows point at.
|
||||
/// Unlike AddBlog this does not skip "deact" names: a note by a deactivated blog still
|
||||
/// needs an ID, and Blogs is the only place one can live.
|
||||
/// Assigning the ID is bookkeeping, not a content change, so DateModified is not touched.
|
||||
/// MAX(BlogId) + 1 cannot hand out a used ID because trg_Blogs_BlogId_NoDelete stops any
|
||||
/// row that holds one from being deleted.
|
||||
/// </summary>
|
||||
private static void RegisterBlogName(SQLiteConnection connection, SQLiteTransaction? transaction, string blogName)
|
||||
private static void RegisterBlog(SQLiteConnection connection, SQLiteTransaction? transaction, string blogName)
|
||||
{
|
||||
using SQLiteCommand command = new SQLiteCommand("INSERT OR IGNORE INTO BlogNames (BlogName) VALUES (@BlogName)", connection, transaction);
|
||||
command.Parameters.AddWithValue("@BlogName", blogName);
|
||||
command.ExecuteNonQuery();
|
||||
string now = DateTime.Now.ToString("yyyy-MM-dd HH:mm:ss");
|
||||
|
||||
using (SQLiteCommand command = new SQLiteCommand("INSERT OR IGNORE INTO Blogs (BlogName, DateAdded, DateModified, DateCreated) VALUES (@BlogName, @Now, @Now, @Now)", connection, transaction))
|
||||
{
|
||||
command.Parameters.AddWithValue("@BlogName", blogName);
|
||||
command.Parameters.AddWithValue("@Now", now);
|
||||
command.ExecuteNonQuery();
|
||||
}
|
||||
|
||||
using (SQLiteCommand command = new SQLiteCommand("UPDATE Blogs SET BlogId = (SELECT IFNULL(MAX(BlogId), 0) + 1 FROM Blogs) WHERE BlogName = @BlogName AND BlogId IS NULL", connection, transaction))
|
||||
{
|
||||
command.Parameters.AddWithValue("@BlogName", blogName);
|
||||
command.ExecuteNonQuery();
|
||||
}
|
||||
}
|
||||
|
||||
/// <summary>
|
||||
@@ -597,6 +614,7 @@ namespace URLNotesGrabberCORE
|
||||
{
|
||||
DBPath ??= GetDefaultDbPath();
|
||||
postType = PostTypes.Normalize(postType);
|
||||
postDate = PostDates.Normalize(postDate)!;
|
||||
try { AddBlog(blogName, byLikes, DBPath); } catch { }
|
||||
try { UpdatePostSetDate(blogName, postID, postDate, DBPath); } catch { }
|
||||
try { UpdatePost(blogName, postID, reblogURL, postDate, postURL, slug, reblogKey, reblogName, summary, quote, body, tags, link, photoURL, photoCaption, downloadedFiles, audioCaption, question, answer, title, hasImage, byLikes, DBPath, rootBlogName, rootURL, postType); } catch { }
|
||||
@@ -761,7 +779,6 @@ namespace URLNotesGrabberCORE
|
||||
{
|
||||
DBPath ??= GetDefaultDbPath();
|
||||
//try { AddPost(rootBlogName, postID, DBPath); } catch { }
|
||||
try { AddBlog(noteBlogName, false, DBPath); } catch { }
|
||||
|
||||
using SQLiteConnection connection2 = new SQLiteConnection("Data Source=" + DBPath);
|
||||
int rowsInserted = 0;
|
||||
@@ -770,23 +787,24 @@ namespace URLNotesGrabberCORE
|
||||
{
|
||||
connection2.Open();
|
||||
|
||||
// Notes stores integer IDs, so both participants and the type have to exist in
|
||||
// their lookup table before the note can point at them.
|
||||
// Notes stores integer IDs, so both participants need a Blogs row with a BlogId,
|
||||
// and the type a NoteTypes row, before the note can point at them. RegisterBlog
|
||||
// also does what the AddBlog call here used to: register the note's blog.
|
||||
//
|
||||
// All four statements run in one transaction so a crash cannot leave a name or a
|
||||
// All statements run in one transaction so a crash cannot leave a blog or a
|
||||
// type registered with no note. The transaction is committed before the console
|
||||
// output below, which sleeps -- a write lock must not be held across that.
|
||||
using (SQLiteTransaction transaction = connection2.BeginTransaction())
|
||||
{
|
||||
RegisterBlogName(connection2, transaction, rootBlogName);
|
||||
RegisterBlogName(connection2, transaction, noteBlogName);
|
||||
RegisterBlog(connection2, transaction, rootBlogName);
|
||||
RegisterBlog(connection2, transaction, noteBlogName);
|
||||
RegisterNoteType(connection2, transaction, type ?? string.Empty);
|
||||
|
||||
// INSERT OR IGNORE, and no IsActive in the column list: re-crawling a note
|
||||
// that was removed elsewhere leaves the existing row -- and its flag -- alone.
|
||||
string sql = "INSERT OR IGNORE INTO Notes (RootBlogId, NoteBlogId, PostID, TimeStamp, TypeId, DatetimeCrawled, DateModified, DateCreated) " +
|
||||
"SELECT (SELECT BlogId FROM BlogNames WHERE BlogName = @rootBlogName), " +
|
||||
" (SELECT BlogId FROM BlogNames WHERE BlogName = @noteBlogName), " +
|
||||
"SELECT (SELECT BlogId FROM Blogs WHERE BlogName = @rootBlogName), " +
|
||||
" (SELECT BlogId FROM Blogs WHERE BlogName = @noteBlogName), " +
|
||||
" @PostID, @TimeStamp, " +
|
||||
" (SELECT TypeId FROM NoteTypes WHERE Type = @Type), " +
|
||||
" @DatetimeCrawled, @DateModified, @DateCreated";
|
||||
@@ -858,9 +876,12 @@ namespace URLNotesGrabberCORE
|
||||
///
|
||||
/// </summary>
|
||||
/// <param name="withoutNotesOnly"></param>
|
||||
/// <param name="ignoreRefreshCooldown">Drops the age gate on the periodic re-queue branch (--force).</param>
|
||||
/// <param name="fromDate">Lower bound on the *original post's* PostDate for the periodic re-queue branch (--fromDate). Independent of ignoreRefreshCooldown -- applies whether or not --force is also given.</param>
|
||||
/// <param name="toDate">Upper bound on the *original post's* PostDate for the periodic re-queue branch (--toDate). Same independence from ignoreRefreshCooldown as fromDate.</param>
|
||||
/// <param name="DBPath"></param>
|
||||
/// <returns>blogName, postID, lastNoteTimestamp, notesGatheredTimestamp</returns>
|
||||
public static List<Tuple<string, long, long, long>> GetPosts(bool withoutNotesOnly = false, DateTime? beforeDate = null, string? blogName = null, string? DBPath = null)
|
||||
public static List<Tuple<string, long, long, long>> GetPosts(bool withoutNotesOnly = false, DateTime? beforeDate = null, string? blogName = null, bool ignoreRefreshCooldown = false, DateTime? fromDate = null, DateTime? toDate = null, string? DBPath = null)
|
||||
{
|
||||
DBPath ??= GetDefaultDbPath();
|
||||
using SQLiteConnection connection = new SQLiteConnection("Data Source=" + DBPath);
|
||||
@@ -885,16 +906,48 @@ namespace URLNotesGrabberCORE
|
||||
beforeDateFilter = $"WHERE (U.NotesGatheredDateTime < {unixTimestamp} OR U.NotesGatheredDateTime IS NULL)" + Environment.NewLine;
|
||||
}
|
||||
|
||||
// Blogs.IsActive is the crawler's work-selection flag (Rolodex removal sets it to 0)
|
||||
// and is independent of Posts.IsActive/Notes.IsActive -- deactivating a blog does not
|
||||
// touch its posts' own IsActive column. AndIsActive("Posts", ...) above therefore does
|
||||
// not catch a deactivated blog; this join against the source rows is what does, so a
|
||||
// blog taken IsActive = 0 in Blogs stops being re-queued by --collect 1 even if its
|
||||
// posts were never individually marked inactive. Blogs.BlogName is that table's PRIMARY
|
||||
// KEY, so the join rides an index rather than scanning it.
|
||||
//
|
||||
// Same WHERE/AND juggling WhereIsActive does, extended to the optional blog predicate:
|
||||
// either clause may be absent, so the first one present has to open the WHERE.
|
||||
string sourceClause = AndIsActive("Posts", "P", DBPath) + (filterByBlog ? " AND P.BlogName = @blogName" : string.Empty);
|
||||
string sourceClause = AndIsActive("Posts", "P", DBPath) + " AND COALESCE(BL.IsActive, 1) = 1" + (filterByBlog ? " AND P.BlogName = @blogName" : string.Empty);
|
||||
string sourceFilter = sourceClause.Length == 0 ? string.Empty : " WHERE" + sourceClause.Substring(" AND".Length);
|
||||
|
||||
// The zomb-eh branch re-queues that blog's *already collected* posts every 3 days. It needs
|
||||
// no blog-filter handling of its own: it reads PostsWithCount, which the filter has already
|
||||
// scoped, so it contributes its rows when the filter names zomb-eh and nothing otherwise.
|
||||
// That keeps a filtered worklist a strict subset of the unfiltered one -- "--collect 1 X"
|
||||
// returns exactly the rows "--collect 1" would have returned for X.
|
||||
// no blog-filter or IsActive handling of its own: it reads PostsWithCount, which the source
|
||||
// filter above -- Blogs.IsActive included -- has already scoped, so it contributes its rows
|
||||
// only when zomb-eh itself is still IsActive = 1 there. That keeps a filtered worklist a
|
||||
// strict subset of the unfiltered one -- "--collect 1 X" returns exactly the rows
|
||||
// "--collect 1" would have returned for X.
|
||||
//
|
||||
// --force drops the age gate only. NotFound = 0 and the IsActive/blog scoping above still
|
||||
// apply: the flag is "re-collect early", not "collect rows every other path excludes".
|
||||
string refreshCooldownClause = ignoreRefreshCooldown
|
||||
? string.Empty
|
||||
: " AND NotesGatheredDateTime < unixepoch('now', 'localtime', '-3 days')" + Environment.NewLine;
|
||||
|
||||
// --fromDate bounds the *original post's* PostDate, not the re-collect cooldown --
|
||||
// it stacks with refreshCooldownClause instead of replacing it, so it applies the
|
||||
// same way whether or not --force also dropped the cooldown. A NULL PostDate never
|
||||
// satisfies ">=" and is excluded, same as an unfiltered run would still include it
|
||||
// (there's nothing to compare here, so this only narrows, never widens, the result).
|
||||
string fromDateClause = fromDate.HasValue
|
||||
? " AND PostDate >= @fromDate" + Environment.NewLine
|
||||
: string.Empty;
|
||||
|
||||
// --toDate is the same deal, mirrored: stacks alongside fromDateClause/
|
||||
// refreshCooldownClause rather than replacing either, so --fromDate and --toDate
|
||||
// can be given together (or alone) and both hold with or without --force.
|
||||
string toDateClause = toDate.HasValue
|
||||
? " AND PostDate <= @toDate" + Environment.NewLine
|
||||
: string.Empty;
|
||||
|
||||
string refreshBranch =
|
||||
"" + Environment.NewLine +
|
||||
" UNION " + Environment.NewLine +
|
||||
@@ -909,7 +962,9 @@ namespace URLNotesGrabberCORE
|
||||
" FROM PostsWithCount" + Environment.NewLine +
|
||||
" WHERE BlogName = 'zomb-eh'" + Environment.NewLine +
|
||||
" AND NotFound = 0" + Environment.NewLine +
|
||||
" AND NotesGatheredDateTime < unixepoch('now', 'localtime', '-3 days')" + Environment.NewLine;
|
||||
refreshCooldownClause +
|
||||
fromDateClause +
|
||||
toDateClause;
|
||||
|
||||
sql = "WITH PostsWithCount AS" + Environment.NewLine +
|
||||
"(" + Environment.NewLine +
|
||||
@@ -922,7 +977,7 @@ namespace URLNotesGrabberCORE
|
||||
" P.HasNotesGathered," + Environment.NewLine +
|
||||
" P.NotFound," + Environment.NewLine +
|
||||
" P.PostDate" + Environment.NewLine +
|
||||
" FROM Posts P" + sourceFilter + Environment.NewLine +
|
||||
" FROM Posts P LEFT JOIN Blogs BL ON BL.BlogName = P.BlogName" + sourceFilter + Environment.NewLine +
|
||||
")," + Environment.NewLine +
|
||||
"Unioned AS" + Environment.NewLine +
|
||||
"(" + Environment.NewLine +
|
||||
@@ -993,6 +1048,13 @@ namespace URLNotesGrabberCORE
|
||||
if (filterByBlog)
|
||||
command.Parameters.AddWithValue("@blogName", blogName);
|
||||
|
||||
// Only ever referenced by the zomb-eh refresh branch, which only exists when
|
||||
// withoutNotesOnly is true -- harmless to bind unconditionally otherwise.
|
||||
if (fromDate.HasValue)
|
||||
command.Parameters.AddWithValue("@fromDate", fromDate.Value.ToString("yyyy-MM-dd HH:mm:ss"));
|
||||
if (toDate.HasValue)
|
||||
command.Parameters.AddWithValue("@toDate", toDate.Value.ToString("yyyy-MM-dd HH:mm:ss"));
|
||||
|
||||
using (SQLiteDataReader reader = command.ExecuteReader())
|
||||
{
|
||||
while (reader.Read())
|
||||
@@ -1081,11 +1143,11 @@ namespace URLNotesGrabberCORE
|
||||
{
|
||||
connection.Open();
|
||||
|
||||
string sql = "SELECT DISTINCT BN.BlogName as blogName, N.PostID" +
|
||||
string sql = "SELECT DISTINCT RB.BlogName as blogName, N.PostID" +
|
||||
" FROM Notes N" +
|
||||
" INNER JOIN BlogNames BN ON BN.BlogId = N.RootBlogId" +
|
||||
" INNER JOIN Blogs RB ON RB.BlogId = N.RootBlogId" +
|
||||
" WHERE N.TypeId = (SELECT TypeId FROM NoteTypes WHERE Type = 'reply')" + AndIsActive("Notes", "N", DBPath) +
|
||||
" ORDER BY BN.BlogName, N.PostID";
|
||||
" ORDER BY RB.BlogName, N.PostID";
|
||||
|
||||
using (SQLiteCommand command = new SQLiteCommand(sql, connection))
|
||||
{
|
||||
@@ -1126,11 +1188,11 @@ namespace URLNotesGrabberCORE
|
||||
connection.Open();
|
||||
|
||||
// Grouped on the integer rather than the name: the group key is what gets sorted,
|
||||
// and BN.BlogName comes along for free off the join.
|
||||
string sql = @"SELECT BN.BlogName as blogName, N.PostID,
|
||||
// and RB.BlogName comes along for free off the join.
|
||||
string sql = @"SELECT RB.BlogName as blogName, N.PostID,
|
||||
MAX(N.TimeStamp) as LatestTimestamp
|
||||
FROM Notes N
|
||||
INNER JOIN BlogNames BN ON BN.BlogId = N.RootBlogId
|
||||
INNER JOIN Blogs RB ON RB.BlogId = N.RootBlogId
|
||||
WHERE N.TypeId = (SELECT TypeId FROM NoteTypes WHERE Type = 'reply')
|
||||
AND (N.replyText IS NULL OR N.replyText = '' OR N.replyText = '.')" + AndIsActive("Notes", "N", DBPath) + @"
|
||||
GROUP BY N.RootBlogId, N.PostID
|
||||
@@ -1175,13 +1237,13 @@ namespace URLNotesGrabberCORE
|
||||
{
|
||||
connection.Open();
|
||||
|
||||
// Posts carries only BlogName, so this is the one join to Notes that has to go
|
||||
// through BlogNames -- there is no Posts.BlogId to hop on. The name predicate is
|
||||
// pushed into the 20k-row lookup, which then feeds integers to the Notes key.
|
||||
// Posts carries only BlogName, so the join to Notes goes through Blogs -- there is
|
||||
// no Posts.BlogId to hop on. Each post's name is a primary-key probe on Blogs,
|
||||
// which then feeds an integer to the Notes key.
|
||||
string sql = @"SELECT DISTINCT P.BlogName, P.PostID, MAX(N.TimeStamp) as LatestTimestamp
|
||||
FROM Posts P
|
||||
INNER JOIN BlogNames RBN ON RBN.BlogName = P.BlogName
|
||||
INNER JOIN Notes N ON N.RootBlogId = RBN.BlogId AND N.PostID = P.PostID
|
||||
INNER JOIN Blogs RB ON RB.BlogName = P.BlogName
|
||||
INNER JOIN Notes N ON N.RootBlogId = RB.BlogId AND N.PostID = P.PostID
|
||||
WHERE P.NotFound = 0
|
||||
AND N.TypeId = (SELECT TypeId FROM NoteTypes WHERE Type = 'reply')
|
||||
AND (N.replyText IS NULL OR N.replyText = '' OR N.replyText = '.')" + AndIsActive("Posts", "P", DBPath) + AndIsActive("Notes", "N", DBPath) + @"
|
||||
@@ -1370,8 +1432,7 @@ namespace URLNotesGrabberCORE
|
||||
try
|
||||
{
|
||||
connection.Open();
|
||||
// Blogs is reached in one integer hop off Blogs.BlogId, not through BlogNames --
|
||||
// that would add a hop and end in the text comparison the migration removed.
|
||||
// Blogs is reached in one integer hop off Blogs.BlogId.
|
||||
// The negated form is only correct because Notes.TypeId is NOT NULL.
|
||||
string sql = "";
|
||||
if (reblogsOnly)
|
||||
@@ -1658,8 +1719,8 @@ namespace URLNotesGrabberCORE
|
||||
connection.Open();
|
||||
|
||||
string sql = "UPDATE Notes SET TimeStamp = @timestamp, DateModified = @dateModified " +
|
||||
"WHERE RootBlogId = (SELECT BlogId FROM BlogNames WHERE BlogName = @rootBlogName) " +
|
||||
"AND NoteBlogId = (SELECT BlogId FROM BlogNames WHERE BlogName = @noteBlogName) " +
|
||||
"WHERE RootBlogId = (SELECT BlogId FROM Blogs WHERE BlogName = @rootBlogName) " +
|
||||
"AND NoteBlogId = (SELECT BlogId FROM Blogs WHERE BlogName = @noteBlogName) " +
|
||||
"AND PostID = @postID AND IFNULL(TimeStamp, 0) <> @timestamp";
|
||||
using (SQLiteCommand command = new SQLiteCommand(sql, connection))
|
||||
{
|
||||
@@ -1973,7 +2034,7 @@ namespace URLNotesGrabberCORE
|
||||
// Only fan out to rows that match the SELECT criteria in GetRepliesWithFilledText (NULL/empty/legacy-'.'). Never overwrite '?' (confirmed-empty) or already-fetched text.
|
||||
// The ABS() term cannot use an index on TimeStamp, before or after the integer schema; the NoteBlogId probe is what keeps this off a full scan.
|
||||
string sql = "UPDATE Notes SET replyText = @replyText, DateModified = @dateModified " +
|
||||
"WHERE NoteBlogId = (SELECT BlogId FROM BlogNames WHERE BlogName = @noteBlogName) " +
|
||||
"WHERE NoteBlogId = (SELECT BlogId FROM Blogs WHERE BlogName = @noteBlogName) " +
|
||||
"AND ABS(TimeStamp - @TimeStamp) <= 5 " +
|
||||
"AND TypeId = (SELECT TypeId FROM NoteTypes WHERE Type = 'reply') " +
|
||||
"AND (replyText IS NULL OR replyText = '' OR replyText = '.') " +
|
||||
@@ -2019,7 +2080,7 @@ namespace URLNotesGrabberCORE
|
||||
connection.Open();
|
||||
|
||||
string sql = "UPDATE Notes SET replyText = @replyText, DateModified = @dateModified " +
|
||||
"WHERE RootBlogId = (SELECT BlogId FROM BlogNames WHERE BlogName = @rootBlogName) " +
|
||||
"WHERE RootBlogId = (SELECT BlogId FROM Blogs WHERE BlogName = @rootBlogName) " +
|
||||
"AND PostID = @PostID " +
|
||||
"AND TypeId = (SELECT TypeId FROM NoteTypes WHERE Type = 'reply') " +
|
||||
"AND IFNULL(replyText, '.') <> @replyText";
|
||||
@@ -2205,6 +2266,7 @@ namespace URLNotesGrabberCORE
|
||||
// column. PostType is used as an output filename, so this is the invariant that keeps
|
||||
// a stray value from becoming a stray file.
|
||||
postType = PostTypes.Normalize(postType);
|
||||
postDate = PostDates.Normalize(postDate);
|
||||
try { AddBlog(blogName, false, DBPath); } catch { }
|
||||
|
||||
SQLiteConnection connection;
|
||||
@@ -2576,10 +2638,11 @@ namespace URLNotesGrabberCORE
|
||||
if (string.IsNullOrWhiteSpace(kvp.Value)) continue;
|
||||
string? column = MapPrefixToColumn(kvp.Key);
|
||||
if (column == null) continue;
|
||||
string value = column == "PostDate" ? PostDates.Normalize(kvp.Value)! : kvp.Value;
|
||||
string paramName = "@p" + parameters.Count;
|
||||
setClauses.Add($"{column} = {paramName}");
|
||||
changedClauses.Add($"IFNULL({column}, '') <> {paramName}");
|
||||
parameters.Add((paramName, kvp.Value));
|
||||
parameters.Add((paramName, value));
|
||||
}
|
||||
|
||||
if (setClauses.Count == 0) return false;
|
||||
|
||||
@@ -0,0 +1,27 @@
|
||||
using System;
|
||||
using System.Globalization;
|
||||
|
||||
namespace URLNotesGrabberCORE
|
||||
{
|
||||
/// <summary>
|
||||
/// The single format for Posts.PostDate: "yyyy-MM-dd HH:mm:ss GMT", which is what the
|
||||
/// Tumblr API sends and what nearly every row holds. Text-file exports can carry the
|
||||
/// RFC 1123 form ("Fri, 14 Feb 2025 15:20:09 GMT") instead, which as text sorts on its
|
||||
/// weekday name and falls outside every --fromDate/--toDate range comparison.
|
||||
///
|
||||
/// Every path that writes PostDate routes through <see cref="Normalize"/>. Only the RFC 1123
|
||||
/// form is rewritten; anything else, including the "." no-change sentinel, passes through.
|
||||
/// </summary>
|
||||
public static class PostDates
|
||||
{
|
||||
public static string? Normalize(string? value)
|
||||
{
|
||||
if (string.IsNullOrWhiteSpace(value)) return value;
|
||||
string trimmed = value.Trim();
|
||||
if (DateTime.TryParseExact(trimmed, "r", CultureInfo.InvariantCulture,
|
||||
DateTimeStyles.AdjustToUniversal | DateTimeStyles.AssumeUniversal, out DateTime parsed))
|
||||
return parsed.ToString("yyyy-MM-dd HH:mm:ss", CultureInfo.InvariantCulture) + " GMT";
|
||||
return value;
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -44,6 +44,8 @@ namespace URLNotesGrabberCORE
|
||||
bool apiExplicitlySet = false;
|
||||
string startFromBlogName = string.Empty;
|
||||
bool forceIgnoreCooldown = false;
|
||||
DateTime? fromDate = null;
|
||||
DateTime? toDate = null;
|
||||
List<string> filteredArgs = new List<string>();
|
||||
for (int i = 0; i < args.Length; i++)
|
||||
{
|
||||
@@ -60,6 +62,34 @@ namespace URLNotesGrabberCORE
|
||||
continue;
|
||||
}
|
||||
|
||||
if (string.Equals(args[i], "--fromDate", StringComparison.OrdinalIgnoreCase))
|
||||
{
|
||||
if (i + 1 < args.Length && DateTime.TryParse(args[i + 1], out DateTime parsedFromDate))
|
||||
{
|
||||
fromDate = parsedFromDate;
|
||||
i++;
|
||||
}
|
||||
else
|
||||
{
|
||||
Console.WriteLine("--Missing or unparseable date after --fromDate. Ignoring.--");
|
||||
}
|
||||
continue;
|
||||
}
|
||||
|
||||
if (string.Equals(args[i], "--toDate", StringComparison.OrdinalIgnoreCase))
|
||||
{
|
||||
if (i + 1 < args.Length && DateTime.TryParse(args[i + 1], out DateTime parsedToDate))
|
||||
{
|
||||
toDate = parsedToDate;
|
||||
i++;
|
||||
}
|
||||
else
|
||||
{
|
||||
Console.WriteLine("--Missing or unparseable date after --toDate. Ignoring.--");
|
||||
}
|
||||
continue;
|
||||
}
|
||||
|
||||
if (string.Equals(args[i], "--api3", StringComparison.OrdinalIgnoreCase))
|
||||
{
|
||||
apiSectionName = "TumblrApi3";
|
||||
@@ -322,7 +352,22 @@ namespace URLNotesGrabberCORE
|
||||
managedCollectRun = true;
|
||||
}
|
||||
|
||||
exitCode = CollectNotes(settings.GetValue<string>("PathOutput"), withoutNotesOnly, beforeDate, managedCollectRun, collectBlogName).GetAwaiter().GetResult();
|
||||
if (forceIgnoreCooldown)
|
||||
Console.WriteLine(withoutNotesOnly
|
||||
? "--force: ignoring the periodic re-collect cooldown; already-collected posts in scope are re-queued now"
|
||||
: "--force: no effect in mode 0 - a full re-check already re-collects every post");
|
||||
|
||||
if (fromDate.HasValue)
|
||||
Console.WriteLine(withoutNotesOnly
|
||||
? $"--fromDate: only re-queuing already-collected posts originally posted on/after {fromDate.Value} (applies with or without --force)"
|
||||
: "--fromDate: no effect in mode 0 - it only bounds the periodic re-queue branch");
|
||||
|
||||
if (toDate.HasValue)
|
||||
Console.WriteLine(withoutNotesOnly
|
||||
? $"--toDate: only re-queuing already-collected posts originally posted on/before {toDate.Value} (applies with or without --force)"
|
||||
: "--toDate: no effect in mode 0 - it only bounds the periodic re-queue branch");
|
||||
|
||||
exitCode = CollectNotes(settings.GetValue<string>("PathOutput"), withoutNotesOnly, beforeDate, managedCollectRun, collectBlogName, forceIgnoreCooldown, fromDate, toDate).GetAwaiter().GetResult();
|
||||
break;
|
||||
|
||||
case "--blogsR": //collect notes from all posts
|
||||
@@ -454,7 +499,7 @@ namespace URLNotesGrabberCORE
|
||||
|
||||
Console.WriteLine("--blogs\t For each Blog in DB, write blogname to file");
|
||||
|
||||
Console.WriteLine("--collect [0|1] [datetime] [blogname]\t Collect Notes from API. 1=only posts without notes. 0=full re-check of all posts: a single resumable pass (interrupt & relaunch to resume; stops when complete, retrigger for a new pass). Optional datetime overrides the cutoff and runs as a one-off (bypasses resume tracking). Optional blogname restricts the run to that blog (exact, case-sensitive match) and also runs as a one-off; e.g. \"--collect 1 zomb-eh\". datetime and blogname may be given in either order - use --blog=name if a blog name would otherwise parse as a date.");
|
||||
Console.WriteLine("--collect [0|1] [datetime] [blogname]\t Collect Notes from API. 1=only posts without notes. 0=full re-check of all posts: a single resumable pass (interrupt & relaunch to resume; stops when complete, retrigger for a new pass). Optional datetime overrides the cutoff and runs as a one-off (bypasses resume tracking). Optional blogname restricts the run to that blog (exact, case-sensitive match) and also runs as a one-off; e.g. \"--collect 1 zomb-eh\". datetime and blogname may be given in either order - use --blog=name if a blog name would otherwise parse as a date. Add --force to ignore the periodic re-collect cooldown and re-queue already-collected posts immediately (mode 1 only). Add --fromDate <datetime> / --toDate <datetime> to only re-queue already-collected posts originally posted on/after / on/before that date (mode 1 only; either or both may be given; applies with or without --force).");
|
||||
|
||||
Console.WriteLine("--blogsR\t For each Note that is a REPLY, write blogname to file ");
|
||||
|
||||
@@ -466,7 +511,11 @@ namespace URLNotesGrabberCORE
|
||||
|
||||
Console.WriteLine("--likes\t Fetch likes: initial backfill for new blogs, incremental refresh for blogs past cooldown. Optional blog name forces single-blog run.");
|
||||
|
||||
Console.WriteLine("--force\t (with --likes) Ignore cooldown and refresh every fully-backfilled blog");
|
||||
Console.WriteLine("--force\t Ignore refresh cooldowns: with --likes, refresh every fully-backfilled blog; with --collect 1, re-queue already-collected posts without waiting out their cooldown");
|
||||
|
||||
Console.WriteLine("--fromDate <datetime>\t With --collect 1, only re-queue already-collected posts originally posted on/after <datetime>. Independent of --force - applies whether or not the cooldown is also bypassed.");
|
||||
|
||||
Console.WriteLine("--toDate <datetime>\t With --collect 1, only re-queue already-collected posts originally posted on/before <datetime>. Independent of --force; may be combined with --fromDate for a range.");
|
||||
|
||||
Console.WriteLine("--urldump\t Scan all posts' text columns and extract suspected URLs to configured file");
|
||||
|
||||
@@ -1345,15 +1394,17 @@ if (shouldInsert)
|
||||
// blipping on one post. Past this, skipping post-by-post would just hammer a closed door.
|
||||
const int MaxConsecutiveTransient = 10;
|
||||
|
||||
static async Task<int> CollectNotes(string outPath, bool withoutNotesOnly = true, DateTime? beforeDate = null, bool managedRun = false, string? blogName = null)
|
||||
static async Task<int> CollectNotes(string outPath, bool withoutNotesOnly = true, DateTime? beforeDate = null, bool managedRun = false, string? blogName = null, bool ignoreRefreshCooldown = false, DateTime? fromDate = null, DateTime? toDate = null)
|
||||
{
|
||||
List<Tuple<string, long, long, long>> posts = DataAccess.GetPosts(withoutNotesOnly, beforeDate, blogName);
|
||||
List<Tuple<string, long, long, long>> posts = DataAccess.GetPosts(withoutNotesOnly, beforeDate, blogName, ignoreRefreshCooldown, fromDate, toDate);
|
||||
|
||||
if (posts.Count == 0 && !string.IsNullOrWhiteSpace(blogName))
|
||||
{
|
||||
// BlogName is matched exactly, so a typo or a case mismatch looks identical to "nothing
|
||||
// left to collect". Say so rather than reporting a silent, instant success.
|
||||
Console.WriteLine($"No posts to collect for blog '{blogName}'. Either it is fully collected, or the name does not match a stored blog (the match is case-sensitive).");
|
||||
if (withoutNotesOnly && !ignoreRefreshCooldown)
|
||||
Console.WriteLine("Already-collected posts are re-queued only once their cooldown elapses; add --force to re-collect them now.");
|
||||
return 0;
|
||||
}
|
||||
|
||||
@@ -1451,7 +1502,7 @@ if (shouldInsert)
|
||||
}
|
||||
|
||||
// Re-fetch the updated list after processing the current post
|
||||
posts = DataAccess.GetPosts(withoutNotesOnly, beforeDate, blogName);
|
||||
posts = DataAccess.GetPosts(withoutNotesOnly, beforeDate, blogName, ignoreRefreshCooldown, fromDate, toDate);
|
||||
}
|
||||
}
|
||||
|
||||
|
||||
+116
-94
@@ -24,21 +24,43 @@ Everything below was read out of the live file, not inferred from code. Counts a
|
||||
> Applied by `../normalize-notes.sql`, which took the file from 207 MB to 148 MB. An
|
||||
> earlier change the same day (`../shrink-db.sql`) took it from 267 MB to 207 MB.
|
||||
|
||||
> ### ⚠ Breaking change, 2026-09-28: `BlogNames` is gone; `Blogs.BlogId` is the only ID authority
|
||||
>
|
||||
> The IDs in `Notes` used to live in a `BlogNames` table, with a copy in `Blogs.BlogId`.
|
||||
> Nothing kept the copy current, so by 2026-09-28 12,238 blogs first seen after the
|
||||
> migration had `Blogs.BlogId = NULL`. Every `Notes`-to-`Blogs` join on `BlogId` silently
|
||||
> skipped them and their 23,148 notes, which kept them out of `GetBlogs`.
|
||||
>
|
||||
> `../retire-blognames.sql` fixed this by giving every note participant a `Blogs` row,
|
||||
> backfilling the IDs (none renumbered), making `ix_Blogs_BlogId` unique, and **dropping
|
||||
> `BlogNames`**. There is no compatibility view: any query naming it fails with
|
||||
> `no such table: BlogNames`. Two triggers now protect the IDs.
|
||||
>
|
||||
> **Porting an app:** replace `BlogNames` with `Blogs` everywhere. The columns you used,
|
||||
> `BlogId` and `BlogName`, exist there with the same meaning. A name lookup
|
||||
> (`SELECT BlogId FROM Blogs WHERE BlogName = ?`) is a primary-key probe, and an ID
|
||||
> lookup or join (`JOIN Blogs b ON b.BlogId = n.NoteBlogId`) uses the unique
|
||||
> `ix_Blogs_BlogId`. Every ID in `Notes` resolves to exactly one `Blogs` row. `Blogs.BlogId`
|
||||
> is **no longer** a stale copy, so any code or docs that distrust it can drop that
|
||||
> caveat. Never write `BlogId` or `BlogName` on a row that has an ID, and never delete such
|
||||
> a row: the triggers reject all three. See [`Blogs`](#blogs).
|
||||
|
||||
---
|
||||
|
||||
## The three content tables
|
||||
|
||||
| Table | Rows | What it is |
|
||||
|---|--:|---|
|
||||
| `Blogs` | 188,620 | The crawl registry — one row per known blog, plus crawl-state flags |
|
||||
| `Blogs` | 198,560 | The crawl registry, one row per known blog plus crawl-state flags. Also the ID authority for blogs in `Notes` |
|
||||
| `Posts` | 22,468 | Stored post content. Only 3,867 blogs actually have any |
|
||||
| `Notes` | 1,182,333 | The engagement graph: `NoteBlogId` acted on `(RootBlogId, PostID)` |
|
||||
| `Notes` | 1,234,830 | The engagement graph: `NoteBlogId` acted on `(RootBlogId, PostID)` |
|
||||
|
||||
…supported by two lookup tables that exist only to keep `Notes` small:
|
||||
(`Blogs` and `Notes` counts as of 2026-09-28; the rest as of 2026-08-07.)
|
||||
|
||||
…supported by one lookup table that exists only to keep `Notes` small:
|
||||
|
||||
| Table | Rows | What it is |
|
||||
|---|--:|---|
|
||||
| `BlogNames` | 20,430 | `BlogId` ⇄ `BlogName`. The ID authority for everything in `Notes` |
|
||||
| `NoteTypes` | 5 | `TypeId` ⇄ `Type`. `like`, `reblog`, `reply`, `posted`, `post_attribution` |
|
||||
|
||||
The engagement graph is the interesting part. 20,311 distinct blogs appear as engagers —
|
||||
@@ -66,27 +88,47 @@ CREATE TABLE "Blogs" (
|
||||
PRIMARY KEY("BlogName")
|
||||
);
|
||||
|
||||
CREATE INDEX ix_Blogs_BlogId ON Blogs (BlogId);
|
||||
CREATE UNIQUE INDEX ix_Blogs_BlogId ON Blogs (BlogId);
|
||||
|
||||
CREATE TRIGGER trg_Blogs_BlogId_NoDelete -- no DELETE of a row that has a BlogId
|
||||
CREATE TRIGGER trg_Blogs_BlogId_Immutable -- no change to its BlogId or BlogName
|
||||
```
|
||||
|
||||
`BlogName` is the primary key, so it is the only indexed way in by name. There is no index
|
||||
on any flag or date — filtering or sorting on those scans all 188k rows, which is
|
||||
on any flag or date. Filtering or sorting on those scans the whole table, which is
|
||||
affordable here and is not on `Notes`.
|
||||
|
||||
**`BlogId` is new as of 2026-08-07 and is the join key to `Notes`.** It exists so that
|
||||
`Notes` can reach `Blogs` in a single integer hop rather than going through `BlogNames`
|
||||
and ending in a text comparison:
|
||||
**`BlogId` is the ID that `Notes.RootBlogId` and `Notes.NoteBlogId` store, and `Blogs` is
|
||||
the only place it lives** (since 2026-09-28; see the banner at the top). The join to
|
||||
`Notes` is one integer hop on the unique index:
|
||||
|
||||
```sql
|
||||
-- what you want
|
||||
FROM Blogs B JOIN Notes N ON N.NoteBlogId = B.BlogId
|
||||
|
||||
-- not this
|
||||
FROM Blogs B JOIN BlogNames BN ON BN.BlogName = B.BlogName
|
||||
JOIN Notes N ON N.NoteBlogId = BN.BlogId
|
||||
```
|
||||
|
||||
**`BlogId` is NULL on 168,202 of 188,620 rows** — every blog that has never appeared in a
|
||||
**Every blog that appears in `Notes` has a `Blogs` row with a `BlogId`.** `AddNote`
|
||||
guarantees it through `RegisterBlog`, which runs in the note's own transaction:
|
||||
|
||||
```sql
|
||||
INSERT OR IGNORE INTO Blogs (BlogName, DateAdded, DateModified, DateCreated)
|
||||
VALUES (@name, @now, @now, @now);
|
||||
UPDATE Blogs SET BlogId = (SELECT IFNULL(MAX(BlogId), 0) + 1 FROM Blogs)
|
||||
WHERE BlogName = @name AND BlogId IS NULL;
|
||||
```
|
||||
|
||||
- Unlike `AddBlog`, this does **not** skip names containing `deact`. A note by a
|
||||
deactivated blog still needs an ID, so such blogs now get registry rows too, with the
|
||||
usual defaults (`HasBeenOutput = 0`, `IsActive` left at its default).
|
||||
- Assigning a `BlogId` is bookkeeping, so it **does not move `DateModified`**.
|
||||
- `MAX(BlogId) + 1` is safe only because an ID can never be freed. The two triggers see
|
||||
to that: deleting a row that has a `BlogId`, or changing its `BlogId` or `BlogName`,
|
||||
aborts. Remove a blog with `IsActive = 0` instead. A blog renamed upstream gets a new
|
||||
row. Rows with no `BlogId` can still be deleted or renamed freely.
|
||||
- `INSERT OR REPLACE` on `Blogs` gets around the delete trigger (SQLite does not fire
|
||||
delete triggers for REPLACE unless `recursive_triggers` is on), and it would wipe the
|
||||
`BlogId`. It was already forbidden because it resets `IsActive`. Do not use it.
|
||||
|
||||
**`BlogId` is NULL on 165,887 of 198,560 rows**, every blog that has never appeared in a
|
||||
note. That is the large majority, and it is not an error: the registry is far bigger than
|
||||
the engagement graph. An inner join on `BlogId` therefore silently drops those blogs,
|
||||
which is usually what you want for engagement queries and is wrong for registry listings.
|
||||
@@ -154,6 +196,10 @@ Notable:
|
||||
on every row, and something has since started writing it. Anything that treated it as
|
||||
permanently unset, or derived the type from post content instead, should be re-examined
|
||||
against the live data. Rolodex still derives it.
|
||||
- **`PostDate` is `yyyy-MM-dd HH:mm:ss GMT`** — UTC, as the Tumblr API sends it, unlike
|
||||
the local-time `DateCreated`/`DateModified`. Text-file exports may carry RFC 1123
|
||||
(`Fri, 14 Feb 2025 15:20:09 GMT`); every write path runs `PostDates.Normalize` to
|
||||
convert it, and `../normalize-postdate.sql` fixed the 4 rows written before that.
|
||||
- `HasImage = 1` on 14,026 rows. It records that the post *had* a picture, not that a
|
||||
usable URL was kept, so it is not a reliable predictor that anything will render.
|
||||
- `PhotoURL` is largely unused; in practice the image markup lives inside `Body`.
|
||||
@@ -183,8 +229,7 @@ CREATE INDEX ix_Notes_NoteBlogId ON Notes (NoteBlogId);
|
||||
|
||||
**Integer IDs since 2026-08-07 — this is the breaking change.** `RootBlogName`,
|
||||
`NoteBlogName` and `Type` are gone, replaced by `RootBlogId`, `NoteBlogId` and `TypeId`.
|
||||
Resolve them through [`BlogNames`](#blognames) and [`NoteTypes`](#notetypes), or join
|
||||
straight to `Blogs` on `BlogId`. The old names were text repeated on 1.18 million rows,
|
||||
Resolve blog IDs through `Blogs.BlogId` and types through [`NoteTypes`](#notetypes). The old names were text repeated on 1.18 million rows,
|
||||
in the table *and* in every index over it; the swap took the file from 207 MB to 148 MB.
|
||||
|
||||
The **primary key column order is deliberately unchanged**, so the leading-prefix access
|
||||
@@ -247,43 +292,21 @@ At 1.18M rows this is the table that dictates how the whole database has to be q
|
||||
- `replyText` is `'.'` on 1,167,464 rows — only `reply` notes carry real text. Those
|
||||
dots are inherited from the old column default; new rows get `NULL` instead.
|
||||
|
||||
**Resolve IDs by filtering the lookup, not by scanning `Notes`.** The lookup tables are
|
||||
tiny and uniquely indexed, so pushing a name predicate into them costs nothing and lets
|
||||
the `Notes` index do the work:
|
||||
**Resolve IDs by filtering `Blogs`, not by scanning `Notes`.** A name predicate on `Blogs`
|
||||
is a primary-key probe, so pushing it there costs nothing and lets the `Notes` index do
|
||||
the work:
|
||||
|
||||
```sql
|
||||
-- good: BlogNames resolves the name, then the index is searched
|
||||
-- good: Blogs resolves the name, then the index is searched
|
||||
SELECT * FROM Notes
|
||||
WHERE NoteBlogId = (SELECT BlogId FROM BlogNames WHERE BlogName = ?);
|
||||
WHERE NoteBlogId = (SELECT BlogId FROM Blogs WHERE BlogName = ?);
|
||||
|
||||
-- also good, same plan
|
||||
SELECT n.* FROM Notes n
|
||||
JOIN BlogNames b ON b.BlogId = n.NoteBlogId
|
||||
JOIN Blogs b ON b.BlogId = n.NoteBlogId
|
||||
WHERE b.BlogName = ?;
|
||||
```
|
||||
|
||||
### `BlogNames`
|
||||
|
||||
```sql
|
||||
CREATE TABLE BlogNames (
|
||||
BlogId INTEGER PRIMARY KEY,
|
||||
BlogName TEXT NOT NULL UNIQUE
|
||||
);
|
||||
```
|
||||
|
||||
20,430 rows — every name appearing in `Notes` as either participant, and nothing else.
|
||||
This is the **ID authority**: `Notes.RootBlogId` and `Notes.NoteBlogId` both point here,
|
||||
and `Blogs.BlogId` is a copy of the value for the blogs that have one.
|
||||
|
||||
**12 of these names have no `Blogs` row.** The registry has never been a superset of the
|
||||
engagement graph and still is not, so resolving an ID through `Blogs` rather than
|
||||
`BlogNames` will occasionally find nothing. Use `BlogNames` when you need the name itself
|
||||
and `Blogs` when you need registry columns.
|
||||
|
||||
IDs are assigned by SQLite and are **stable**: they are stored in over a million `Notes`
|
||||
rows. Never renumber them. A blog that is renamed upstream should get a new row, not an
|
||||
edit to an existing one, unless every `Notes` reference is migrated with it.
|
||||
|
||||
### `NoteTypes`
|
||||
|
||||
```sql
|
||||
@@ -313,6 +336,9 @@ code to this table's contents, so prefer the join in anything long-lived.
|
||||
|
||||
## Porting to the integer schema
|
||||
|
||||
> Written for the 2026-08-07 change, and updated for 2026-09-28: wherever this section
|
||||
> once said `BlogNames`, it now says `Blogs`. `BlogNames` no longer exists.
|
||||
|
||||
Everything here was checked against the live 148 MB file. There were 14 affected call
|
||||
sites in `DataAccess.cs` and 16 in `RolodexRepository.cs`. TumblThree needs no changes —
|
||||
its single statement touches `Blogs.IsActive` and `BlogName` only.
|
||||
@@ -329,8 +355,8 @@ the result.
|
||||
|
||||
| Was | Is now | Resolve via |
|
||||
|---|---|---|
|
||||
| `Notes.RootBlogName` | `Notes.RootBlogId` | `BlogNames.BlogId` → `.BlogName` |
|
||||
| `Notes.NoteBlogName` | `Notes.NoteBlogId` | `BlogNames.BlogId` → `.BlogName` |
|
||||
| `Notes.RootBlogName` | `Notes.RootBlogId` | `Blogs.BlogId` → `.BlogName` |
|
||||
| `Notes.NoteBlogName` | `Notes.NoteBlogId` | `Blogs.BlogId` → `.BlogName` |
|
||||
| `Notes.Type` | `Notes.TypeId` | `NoteTypes.TypeId` → `.Type` |
|
||||
| `ix_NoteBlogName01` | `ix_Notes_NoteBlogId` | — |
|
||||
|
||||
@@ -343,14 +369,14 @@ the result.
|
||||
-- was
|
||||
WHERE NoteBlogName = @Name
|
||||
|
||||
-- now, either form; both search ix_Notes_NoteBlogId after a unique-index lookup
|
||||
WHERE NoteBlogId = (SELECT BlogId FROM BlogNames WHERE BlogName = @Name)
|
||||
-- now, either form; both search ix_Notes_NoteBlogId after a primary-key lookup
|
||||
WHERE NoteBlogId = (SELECT BlogId FROM Blogs WHERE BlogName = @Name)
|
||||
-- or
|
||||
JOIN BlogNames b ON b.BlogId = n.NoteBlogId WHERE b.BlogName = @Name
|
||||
JOIN Blogs b ON b.BlogId = n.NoteBlogId WHERE b.BlogName = @Name
|
||||
```
|
||||
|
||||
Measured 73 ms against 63 ms for the old text form on the busiest blog — the extra hop is
|
||||
a unique-index probe on a 20k-row table and does not show.
|
||||
Measured at 73 ms against 63 ms for the old text form on the busiest blog, when the lookup
|
||||
was still `BlogNames`. The extra hop is one index probe and does not show.
|
||||
|
||||
### Joining `Notes` to `Blogs`
|
||||
|
||||
@@ -360,12 +386,17 @@ This is the join to get right; it is the most common shape in both applications.
|
||||
-- was
|
||||
FROM Blogs B INNER JOIN Notes N ON N.NoteBlogName = B.BlogName
|
||||
|
||||
-- now: one integer hop, using the new Blogs.BlogId
|
||||
-- now: one integer hop, using Blogs.BlogId
|
||||
FROM Blogs B INNER JOIN Notes N ON N.NoteBlogId = B.BlogId
|
||||
```
|
||||
|
||||
Do **not** route this through `BlogNames` — that adds a hop and ends in the text
|
||||
comparison the change was meant to remove.
|
||||
Joining `Notes` to `Posts` also goes through `Blogs`, since `Posts` has only a name:
|
||||
|
||||
```sql
|
||||
FROM Posts P
|
||||
JOIN Blogs RB ON RB.BlogName = P.BlogName
|
||||
JOIN Notes N ON N.RootBlogId = RB.BlogId AND N.PostID = P.PostID
|
||||
```
|
||||
|
||||
### Selecting a name back out
|
||||
|
||||
@@ -374,12 +405,12 @@ comparison the change was meant to remove.
|
||||
SELECT NoteBlogName AS blogName, COUNT(*) FROM Notes ... GROUP BY NoteBlogName
|
||||
|
||||
-- now
|
||||
SELECT bn.BlogName AS blogName, COUNT(*)
|
||||
FROM Notes n JOIN BlogNames bn ON bn.BlogId = n.NoteBlogId
|
||||
... GROUP BY bn.BlogName
|
||||
SELECT b.BlogName AS blogName, COUNT(*)
|
||||
FROM Notes n JOIN Blogs b ON b.BlogId = n.NoteBlogId
|
||||
... GROUP BY b.BlogName
|
||||
```
|
||||
|
||||
Group by `n.NoteBlogId` instead of `bn.BlogName` when you only need the name for display —
|
||||
Group by `n.NoteBlogId` instead of `b.BlogName` when you only need the name for display —
|
||||
grouping on the integer is cheaper and the name comes along for free.
|
||||
|
||||
### Filtering by type
|
||||
@@ -403,27 +434,25 @@ because `TypeId` is `NOT NULL`.
|
||||
|
||||
### Inserting a note
|
||||
|
||||
The crawler must ensure both names have IDs first. `INSERT OR IGNORE` on `BlogNames` is
|
||||
the whole of it — no read-back, no round trip, safe to run every time:
|
||||
The crawler must ensure both blogs have IDs first: run the `RegisterBlog` pair shown under
|
||||
[`Blogs`](#blogs) for each name. No read-back, no round trip, and safe to run every time.
|
||||
Then:
|
||||
|
||||
```sql
|
||||
INSERT OR IGNORE INTO BlogNames (BlogName) VALUES (@rootBlogName);
|
||||
INSERT OR IGNORE INTO BlogNames (BlogName) VALUES (@noteBlogName);
|
||||
|
||||
INSERT OR IGNORE INTO Notes
|
||||
(RootBlogId, PostID, NoteBlogId, TimeStamp, TypeId,
|
||||
DatetimeCrawled, DateModified, DateCreated)
|
||||
SELECT (SELECT BlogId FROM BlogNames WHERE BlogName = @rootBlogName),
|
||||
SELECT (SELECT BlogId FROM Blogs WHERE BlogName = @rootBlogName),
|
||||
@PostID,
|
||||
(SELECT BlogId FROM BlogNames WHERE BlogName = @noteBlogName),
|
||||
(SELECT BlogId FROM Blogs WHERE BlogName = @noteBlogName),
|
||||
@TimeStamp,
|
||||
(SELECT TypeId FROM NoteTypes WHERE Type = @Type),
|
||||
@DatetimeCrawled, @DateModified, @DateCreated;
|
||||
```
|
||||
|
||||
Verified: a genuinely new note inserts, and re-running the identical statement inserts 0.
|
||||
Run all three statements in one transaction so a crash cannot leave a name registered
|
||||
with no note.
|
||||
Run the registrations and the insert in one transaction so a crash cannot leave a blog
|
||||
registered with no note.
|
||||
|
||||
**The duplicate-key error message has changed.** `DataAccess.cs` compares against the
|
||||
literal string
|
||||
@@ -448,7 +477,7 @@ on `TimeStamp` either before or after:
|
||||
```sql
|
||||
-- now
|
||||
UPDATE Notes SET replyText = @replyText, DateModified = @dateModified
|
||||
WHERE NoteBlogId = (SELECT BlogId FROM BlogNames WHERE BlogName = @noteBlogName)
|
||||
WHERE NoteBlogId = (SELECT BlogId FROM Blogs WHERE BlogName = @noteBlogName)
|
||||
AND ABS(TimeStamp - @TimeStamp) <= 5
|
||||
AND TypeId = (SELECT TypeId FROM NoteTypes WHERE Type = 'reply')
|
||||
AND (replyText IS NULL OR replyText = '' OR replyText = '.')
|
||||
@@ -458,18 +487,15 @@ UPDATE Notes SET replyText = @replyText, DateModified = @dateModified
|
||||
Rolodex's soft-delete updates need no change beyond the `WHERE` clause — they set
|
||||
`IsActive`, which is untouched.
|
||||
|
||||
### Three traps
|
||||
### Two traps
|
||||
|
||||
**`Blogs.BlogId` is NULL on 168,202 of 188,620 rows.** Any inner join on it silently drops
|
||||
**`Blogs.BlogId` is NULL on 165,887 of 198,560 rows.** Any inner join on it silently drops
|
||||
every blog that has never appeared in a note. Correct for engagement queries; wrong for
|
||||
registry listings, which need a `LEFT JOIN` or no join at all.
|
||||
|
||||
**12 names in `BlogNames` have no `Blogs` row.** Resolving an ID to a name through `Blogs`
|
||||
will occasionally find nothing. Use `BlogNames` for names and `Blogs` for registry columns.
|
||||
|
||||
**IDs are stable and must stay so.** `BlogNames.BlogId` and `NoteTypes.TypeId` are stored
|
||||
in over a million `Notes` rows. Never renumber. A blog renamed upstream gets a new row,
|
||||
not an edited one, unless every `Notes` reference migrates with it.
|
||||
**IDs are stable and must stay so.** `Blogs.BlogId` and `NoteTypes.TypeId` are stored in
|
||||
over a million `Notes` rows. Never renumber. A blog renamed upstream gets a new row, not
|
||||
an edited one. The `Blogs` triggers reject both.
|
||||
|
||||
---
|
||||
|
||||
@@ -477,16 +503,12 @@ not an edited one, unless every `Notes` reference migrates with it.
|
||||
|
||||
There are no foreign keys, and the tables do not perfectly agree:
|
||||
|
||||
- 4 `Posts` rows name a blog with no `Blogs` row.
|
||||
- 12 of the 20,430 names in `BlogNames` have no `Blogs` row.
|
||||
|
||||
So a name appearing in `Notes` or `Posts` is not a guarantee that the registry knows about
|
||||
it. Joins from those tables back to `Blogs` should tolerate a miss.
|
||||
|
||||
The integer schema does not fix this and was not meant to. `BlogNames` is deliberately
|
||||
built from `Notes` rather than from `Blogs`, precisely so that the 12 unregistered
|
||||
engagers keep their IDs and their rows. Had it been built from the registry, those notes
|
||||
would have been dropped by the migration's inner joins.
|
||||
- 4 `Posts` rows name a blog with no `Blogs` row, so joins from `Posts` back to `Blogs`
|
||||
should tolerate a miss.
|
||||
- `Notes` is covered: every `RootBlogId` and `NoteBlogId` resolves to a `Blogs` row.
|
||||
`retire-blognames.sql` checked this before committing, and `RegisterBlog` keeps it true.
|
||||
Before 2026-09-28, 12 to 17 note participants had no registry row. They now have stub
|
||||
rows.
|
||||
|
||||
---
|
||||
|
||||
@@ -538,11 +560,10 @@ Crawler bookkeeping. Rolodex ignores all of these.
|
||||
`DataAccess.cs` joins on it to decide what to collect:
|
||||
|
||||
```sql
|
||||
-- shape only; the ported GetBlogs joins Blogs directly on BlogId and needs no BlogNames hop
|
||||
SELECT bn.BlogName, count(*)
|
||||
-- shape only
|
||||
SELECT b.BlogName, count(*)
|
||||
FROM Notes n
|
||||
JOIN Blogs b ON b.BlogId = n.NoteBlogId
|
||||
JOIN BlogNames bn ON bn.BlogId = n.NoteBlogId
|
||||
JOIN Blogs b ON b.BlogId = n.NoteBlogId
|
||||
WHERE b.IsActive = @isActive AND ...
|
||||
```
|
||||
|
||||
@@ -640,7 +661,7 @@ handled:
|
||||
SELECT 'Blogs', COUNT(*) FROM Blogs
|
||||
UNION ALL SELECT 'Posts', COUNT(*) FROM Posts
|
||||
UNION ALL SELECT 'Notes', COUNT(*) FROM Notes
|
||||
UNION ALL SELECT 'BlogNames', COUNT(*) FROM BlogNames;
|
||||
UNION ALL SELECT 'Blogs with a BlogId', COUNT(*) FROM Blogs WHERE BlogId IS NOT NULL;
|
||||
|
||||
-- note type mix (joins NoteTypes; Notes.Type no longer exists)
|
||||
SELECT t.Type, COUNT(*)
|
||||
@@ -664,8 +685,9 @@ SELECT COUNT(*) FROM (
|
||||
SELECT COUNT(*) FROM Posts p
|
||||
WHERE NOT EXISTS (SELECT 1 FROM Blogs b WHERE b.BlogName = p.BlogName);
|
||||
|
||||
SELECT COUNT(*) FROM BlogNames bn
|
||||
WHERE NOT EXISTS (SELECT 1 FROM Blogs b WHERE b.BlogName = bn.BlogName);
|
||||
-- note participants with no Blogs.BlogId (expect 0; anything else is the pre-2026-09-28 drift)
|
||||
SELECT COUNT(*) FROM (SELECT DISTINCT NoteBlogId AS Id FROM Notes) n
|
||||
WHERE NOT EXISTS (SELECT 1 FROM Blogs b WHERE b.BlogId = n.Id);
|
||||
|
||||
-- space by object, to see where the file actually goes
|
||||
SELECT name, SUM(pgsize)/1024/1024 AS mb
|
||||
|
||||
@@ -2,6 +2,10 @@
|
||||
-- Replaces the repeated blog-name and type TEXT in Notes with integer IDs.
|
||||
-- Reduces TL.db from ~207 MB to ~148 MB (-29%).
|
||||
--
|
||||
-- SUPERSEDED IN PART, 2026-09-28: the BlogNames table this creates is no longer the
|
||||
-- ID authority. Run retire-blognames.sql straight after this one; it moves the IDs
|
||||
-- into Blogs.BlogId and drops BlogNames. The current app code assumes both have run.
|
||||
--
|
||||
-- THIS IS A BREAKING SCHEMA CHANGE. There is no compatibility layer. Every
|
||||
-- query in URLNotesGrabberCORE and Rolodex that names Notes.RootBlogName,
|
||||
-- Notes.NoteBlogName or Notes.Type stops working the moment this runs, and
|
||||
|
||||
@@ -0,0 +1,31 @@
|
||||
-- normalize-postdate.sql
|
||||
-- Rewrites Posts.PostDate values held in RFC 1123 form ("Fri, 14 Feb 2025 15:20:09 GMT")
|
||||
-- into the column's canonical "yyyy-MM-dd HH:mm:ss GMT" (the Tumblr API's own format).
|
||||
--
|
||||
-- As of 2026-09-21 this matched 4 rows, all zombaee, from one text-file import. As text
|
||||
-- they sort on the weekday name and never satisfy --fromDate / --toDate comparisons.
|
||||
-- New writes are normalized in code by PostDates.Normalize, so this is a one-off.
|
||||
--
|
||||
-- Only PostDate changes. DateModified is left alone: the post content did not change.
|
||||
--
|
||||
-- HOW TO RUN: back up TL.db, then from the URLNotesGrabberCORE project folder:
|
||||
-- sqlite3 TL.db < ../normalize-postdate.sql
|
||||
|
||||
SELECT 'before', COUNT(*) FROM Posts WHERE PostDate LIKE '___, __ ___ ____ __:__:__ GMT';
|
||||
|
||||
BEGIN;
|
||||
UPDATE Posts
|
||||
SET PostDate = substr(PostDate, 13, 4) || '-' ||
|
||||
printf('%02d', (instr('JanFebMarAprMayJunJulAugSepOctNovDec', substr(PostDate, 9, 3)) + 2) / 3) || '-' ||
|
||||
substr(PostDate, 6, 2) || ' ' ||
|
||||
substr(PostDate, 18)
|
||||
WHERE PostDate LIKE '___, __ ___ ____ __:__:__ GMT'
|
||||
AND instr('JanFebMarAprMayJunJulAugSepOctNovDec', substr(PostDate, 9, 3)) % 3 = 1;
|
||||
COMMIT;
|
||||
|
||||
-- VERIFY: expect 0, then a single shape '9999-99-99 99:99:99 GMT' (plus any NULL/blank)
|
||||
SELECT 'after', COUNT(*) FROM Posts WHERE PostDate LIKE '___, %';
|
||||
SELECT CASE WHEN PostDate GLOB '[0-9][0-9][0-9][0-9]-[0-9][0-9]-[0-9][0-9] [0-9][0-9]:[0-9][0-9]:[0-9][0-9] GMT'
|
||||
THEN 'yyyy-MM-dd HH:mm:ss GMT' ELSE IFNULL(PostDate, '(null)') END AS shape,
|
||||
COUNT(*)
|
||||
FROM Posts GROUP BY 1 ORDER BY 2 DESC;
|
||||
@@ -0,0 +1,143 @@
|
||||
-- retire-blognames.sql
|
||||
-- Makes Blogs.BlogId the only ID authority for Notes and retires the BlogNames table.
|
||||
--
|
||||
-- WHY: normalize-notes.sql (2026-08-07) put the IDs in BlogNames and copied them into
|
||||
-- Blogs.BlogId once. Nothing kept the copy current: by 2026-09-28, 12,238 blogs first seen
|
||||
-- in a note after the migration had a BlogNames ID but Blogs.BlogId = NULL, so every query
|
||||
-- joining Notes to Blogs on BlogId (GetBlogs and friends) silently skipped them -- 23,148
|
||||
-- notes. Two copies of one ID drift; this leaves one.
|
||||
--
|
||||
-- What it does:
|
||||
-- 1. Gives every BlogNames name a Blogs row (17 had none), carrying its ID over.
|
||||
-- 2. Copies the ID onto every Blogs row that is missing it. IDs are never renumbered --
|
||||
-- they are stored in 1.18M Notes rows.
|
||||
-- 3. Proves every Notes ID resolves through Blogs before anything is dropped.
|
||||
-- 4. Makes ix_Blogs_BlogId UNIQUE.
|
||||
-- 5. Drops BlogNames. No compatibility view: any other app that still names it gets
|
||||
-- "no such table: BlogNames" and must port to Blogs.BlogId (see TL.db.md).
|
||||
-- 6. Adds triggers that stop a Blogs row holding a BlogId from being deleted, renamed or
|
||||
-- renumbered -- the guarantees BlogNames gave by never being touched.
|
||||
--
|
||||
-- DateModified is NOT moved: assigning an ID is bookkeeping, not a content change. The 17
|
||||
-- new stub rows get DateAdded/DateModified/DateCreated = now, as AddBlog would give them.
|
||||
--
|
||||
-- Runs after normalize-notes.sql. A backup from before 2026-08-07 needs both, in order.
|
||||
--
|
||||
-- HOW TO RUN:
|
||||
-- 1. Stop every app that uses TL.db. Pause NextCloud sync.
|
||||
-- 2. Back up TL.db: sqlite3 TL.db ".backup 'TL pre-retire-blognames.db'"
|
||||
-- 3. sqlite3 -bail TL.db < retire-blognames.sql
|
||||
-- -bail matters: a failed check aborts before COMMIT and nothing is changed.
|
||||
-- In DB Browser, Execute SQL stops at the first error; then Revert Changes.
|
||||
-- 4. Run the build of URLNotesGrabberCORE that no longer uses BlogNames. An older
|
||||
-- build fails every AddNote with "no such table: BlogNames".
|
||||
|
||||
PRAGMA foreign_keys = off;
|
||||
|
||||
BEGIN;
|
||||
|
||||
-- Every check inserts one count here; the CHECK aborts the script on anything but 0.
|
||||
CREATE TEMP TABLE MustBeZero (Check_ TEXT, n INTEGER CHECK (n = 0));
|
||||
|
||||
--------------------------------------------------------------------------
|
||||
-- STEP 0: the two copies must not disagree anywhere they are both set
|
||||
--------------------------------------------------------------------------
|
||||
INSERT INTO MustBeZero
|
||||
SELECT 'Blogs.BlogId differs from BlogNames', COUNT(*)
|
||||
FROM Blogs b JOIN BlogNames bn ON bn.BlogName = b.BlogName
|
||||
WHERE b.BlogId <> bn.BlogId;
|
||||
|
||||
INSERT INTO MustBeZero
|
||||
SELECT 'Blogs.BlogId unknown to BlogNames', COUNT(*)
|
||||
FROM Blogs b
|
||||
WHERE b.BlogId IS NOT NULL
|
||||
AND NOT EXISTS (SELECT 1 FROM BlogNames bn WHERE bn.BlogId = b.BlogId AND bn.BlogName = b.BlogName);
|
||||
|
||||
--------------------------------------------------------------------------
|
||||
-- STEP 1: a Blogs row for every name Notes points at
|
||||
--------------------------------------------------------------------------
|
||||
-- No IsActive in the column list: it is not ours to write (defaults to live).
|
||||
INSERT INTO Blogs (BlogName, DateAdded, DateModified, DateCreated, BlogId)
|
||||
SELECT bn.BlogName,
|
||||
strftime('%Y-%m-%d %H:%M:%S', 'now', 'localtime'),
|
||||
strftime('%Y-%m-%d %H:%M:%S', 'now', 'localtime'),
|
||||
strftime('%Y-%m-%d %H:%M:%S', 'now', 'localtime'),
|
||||
bn.BlogId
|
||||
FROM BlogNames bn
|
||||
WHERE NOT EXISTS (SELECT 1 FROM Blogs b WHERE b.BlogName = bn.BlogName);
|
||||
|
||||
--------------------------------------------------------------------------
|
||||
-- STEP 2: backfill the IDs Blogs never received
|
||||
--------------------------------------------------------------------------
|
||||
UPDATE Blogs
|
||||
SET BlogId = (SELECT bn.BlogId FROM BlogNames bn WHERE bn.BlogName = Blogs.BlogName)
|
||||
WHERE BlogId IS NULL
|
||||
AND BlogName IN (SELECT BlogName FROM BlogNames);
|
||||
|
||||
--------------------------------------------------------------------------
|
||||
-- STEP 3: prove Blogs now holds exactly what BlogNames held
|
||||
--------------------------------------------------------------------------
|
||||
INSERT INTO MustBeZero
|
||||
SELECT 'BlogNames pair missing from Blogs', COUNT(*)
|
||||
FROM BlogNames bn
|
||||
WHERE NOT EXISTS (SELECT 1 FROM Blogs b WHERE b.BlogId = bn.BlogId AND b.BlogName = bn.BlogName);
|
||||
|
||||
INSERT INTO MustBeZero
|
||||
SELECT 'Blogs IDs vs BlogNames rows', (SELECT COUNT(*) FROM Blogs WHERE BlogId IS NOT NULL) - (SELECT COUNT(*) FROM BlogNames);
|
||||
|
||||
INSERT INTO MustBeZero
|
||||
SELECT 'Notes.RootBlogId unresolved', COUNT(*)
|
||||
FROM (SELECT DISTINCT RootBlogId AS Id FROM Notes) n
|
||||
WHERE NOT EXISTS (SELECT 1 FROM Blogs b WHERE b.BlogId = n.Id);
|
||||
|
||||
INSERT INTO MustBeZero
|
||||
SELECT 'Notes.NoteBlogId unresolved', COUNT(*)
|
||||
FROM (SELECT DISTINCT NoteBlogId AS Id FROM Notes) n
|
||||
WHERE NOT EXISTS (SELECT 1 FROM Blogs b WHERE b.BlogId = n.Id);
|
||||
|
||||
--------------------------------------------------------------------------
|
||||
-- STEP 4: one row per ID
|
||||
--------------------------------------------------------------------------
|
||||
-- UNIQUE still allows the NULLs on the ~168k blogs that have never appeared in a note.
|
||||
DROP INDEX ix_Blogs_BlogId;
|
||||
CREATE UNIQUE INDEX ix_Blogs_BlogId ON Blogs (BlogId);
|
||||
|
||||
--------------------------------------------------------------------------
|
||||
-- STEP 5: BlogNames goes
|
||||
--------------------------------------------------------------------------
|
||||
DROP TABLE BlogNames;
|
||||
|
||||
--------------------------------------------------------------------------
|
||||
-- STEP 6: what BlogNames guaranteed by never being written
|
||||
--------------------------------------------------------------------------
|
||||
-- A deleted row would orphan its notes, and MAX(BlogId) + 1 in RegisterBlog could then
|
||||
-- hand the same ID to a different blog. Remove a blog with IsActive = 0 instead.
|
||||
CREATE TRIGGER trg_Blogs_BlogId_NoDelete
|
||||
BEFORE DELETE ON Blogs
|
||||
WHEN OLD.BlogId IS NOT NULL
|
||||
BEGIN
|
||||
SELECT RAISE(ABORT, 'Blogs row has a BlogId that Notes points at; set IsActive = 0 instead of deleting');
|
||||
END;
|
||||
|
||||
-- A blog renamed upstream is a new blog to Tumblr's API and gets a new row. Editing the name
|
||||
-- in place would re-attribute every note to it; changing the ID would orphan them.
|
||||
CREATE TRIGGER trg_Blogs_BlogId_Immutable
|
||||
BEFORE UPDATE OF BlogId, BlogName ON Blogs
|
||||
WHEN OLD.BlogId IS NOT NULL
|
||||
AND (NEW.BlogId IS NOT OLD.BlogId OR NEW.BlogName IS NOT OLD.BlogName)
|
||||
BEGIN
|
||||
SELECT RAISE(ABORT, 'BlogId and BlogName are fixed once a blog has a BlogId; Notes rows point at it');
|
||||
END;
|
||||
|
||||
DROP TABLE temp.MustBeZero;
|
||||
|
||||
COMMIT;
|
||||
|
||||
--------------------------------------------------------------------------
|
||||
-- VERIFY
|
||||
--------------------------------------------------------------------------
|
||||
-- SELECT COUNT(*) FROM sqlite_master WHERE name = 'BlogNames'; -- expect: 0
|
||||
-- SELECT COUNT(*) FROM Blogs WHERE BlogId IS NOT NULL; -- expect: the old BlogNames row count
|
||||
-- SELECT sql FROM sqlite_master WHERE name = 'ix_Blogs_BlogId'; -- expect: CREATE UNIQUE INDEX
|
||||
-- SELECT name FROM sqlite_master WHERE type = 'trigger'; -- expect: both triggers
|
||||
-- PRAGMA integrity_check; -- expect: ok
|
||||
+32
-12
@@ -82,7 +82,8 @@ WITH expected(tbl, col, alter_stmt) AS (
|
||||
-- Blogs.BlogId (2026-08-07) is the single-hop join key into Notes. Deliberately NOT
|
||||
-- auto-fixable: an added-but-empty BlogId makes every engagement join return zero
|
||||
-- rows silently, which is worse than the hard error a missing column gives.
|
||||
('Blogs','BlogId', 'MANUAL REVIEW - see query 1d: run normalize-notes.sql'),
|
||||
-- Since 2026-09-28 it is the only blog-ID authority (query 1e).
|
||||
('Blogs','BlogId', 'MANUAL REVIEW - see queries 1d/1e: run normalize-notes.sql, then retire-blognames.sql'),
|
||||
|
||||
-- Notes (base columns: manual review if missing)
|
||||
-- Integer IDs since 2026-08-07. RootBlogName/NoteBlogName/Type are GONE, not renamed
|
||||
@@ -101,11 +102,11 @@ WITH expected(tbl, col, alter_stmt) AS (
|
||||
-- placeholder. EnsureReplyTextColumnExists in DataAccess.cs adds it the same way.
|
||||
('Notes','replyText', 'ALTER TABLE Notes ADD COLUMN replyText TEXT;'),
|
||||
|
||||
-- BlogNames / NoteTypes (the lookup tables Notes resolves its IDs through, 2026-08-07).
|
||||
-- Not auto-fixable: an empty BlogNames does not mean "add the table", it means the
|
||||
-- NoteTypes (the lookup table Notes resolves TypeId through, 2026-08-07).
|
||||
-- Not auto-fixable: an empty NoteTypes does not mean "add the table", it means the
|
||||
-- Notes rows have nothing to resolve against. Rebuild with normalize-notes.sql.
|
||||
('BlogNames','BlogId', 'MANUAL REVIEW - see query 1d: run normalize-notes.sql'),
|
||||
('BlogNames','BlogName', 'MANUAL REVIEW - see query 1d: run normalize-notes.sql'),
|
||||
-- BlogNames is not listed: it was dropped on 2026-09-28. Query 1e reports a file
|
||||
-- that still has it.
|
||||
('NoteTypes','TypeId', 'MANUAL REVIEW - see query 1d: run normalize-notes.sql'),
|
||||
('NoteTypes','Type', 'MANUAL REVIEW - see query 1d: run normalize-notes.sql'),
|
||||
|
||||
@@ -123,7 +124,6 @@ actual(tbl, col) AS (
|
||||
SELECT 'Posts', name FROM pragma_table_info('Posts')
|
||||
UNION ALL SELECT 'Blogs', name FROM pragma_table_info('Blogs')
|
||||
UNION ALL SELECT 'Notes', name FROM pragma_table_info('Notes')
|
||||
UNION ALL SELECT 'BlogNames', name FROM pragma_table_info('BlogNames')
|
||||
UNION ALL SELECT 'NoteTypes', name FROM pragma_table_info('NoteTypes')
|
||||
UNION ALL SELECT 'DailyAPICount', name FROM pragma_table_info('DailyAPICount')
|
||||
UNION ALL SELECT 'ApiKeyPoolState', name FROM pragma_table_info('ApiKeyPoolState')
|
||||
@@ -146,7 +146,7 @@ ORDER BY (e.alter_stmt LIKE 'ALTER%') DESC, e.tbl, e.col;
|
||||
-- 1b. MISSING TABLES: expected tables that don't exist at all in this DB.
|
||||
-- Zero rows = good.
|
||||
WITH expected_tables(tbl) AS (
|
||||
VALUES ('Posts'),('Blogs'),('Notes'),('BlogNames'),('NoteTypes'),('DailyAPICount'),
|
||||
VALUES ('Posts'),('Blogs'),('Notes'),('NoteTypes'),('DailyAPICount'),
|
||||
('ApiKeyPoolState'),('ApiKeyPoolMeta')
|
||||
)
|
||||
SELECT et.tbl AS missing_table
|
||||
@@ -180,7 +180,6 @@ WITH expected(tbl, col) AS (
|
||||
('Notes','RootBlogId'),('Notes','PostID'),('Notes','NoteBlogId'),('Notes','TimeStamp'),
|
||||
('Notes','TypeId'),('Notes','DatetimeCrawled'),('Notes','DateModified'),('Notes','DateCreated'),
|
||||
('Notes','replyText'),('Notes','IsActive'),
|
||||
('BlogNames','BlogId'),('BlogNames','BlogName'),
|
||||
('NoteTypes','TypeId'),('NoteTypes','Type'),
|
||||
('DailyAPICount','Date'),('DailyAPICount','APICount'),
|
||||
('ApiKeyPoolState','KeyName'),('ApiKeyPoolState','RetryUntil'),
|
||||
@@ -190,7 +189,6 @@ actual(tbl, col) AS (
|
||||
SELECT 'Posts', name FROM pragma_table_info('Posts')
|
||||
UNION ALL SELECT 'Blogs', name FROM pragma_table_info('Blogs')
|
||||
UNION ALL SELECT 'Notes', name FROM pragma_table_info('Notes')
|
||||
UNION ALL SELECT 'BlogNames', name FROM pragma_table_info('BlogNames')
|
||||
UNION ALL SELECT 'NoteTypes', name FROM pragma_table_info('NoteTypes')
|
||||
UNION ALL SELECT 'DailyAPICount', name FROM pragma_table_info('DailyAPICount')
|
||||
UNION ALL SELECT 'ApiKeyPoolState', name FROM pragma_table_info('ApiKeyPoolState')
|
||||
@@ -209,7 +207,7 @@ ORDER BY a.tbl, a.col;
|
||||
--
|
||||
-- This is the one failure SECTION 2 cannot fix. Notes.RootBlogName /
|
||||
-- NoteBlogName / Type were replaced by RootBlogId / NoteBlogId / TypeId
|
||||
-- resolving through BlogNames and NoteTypes -- a data migration, not an
|
||||
-- resolving through (then) BlogNames and NoteTypes -- a data migration, not an
|
||||
-- ADD COLUMN. There is no compatibility view, so the current code fails
|
||||
-- outright ("no such column: RootBlogId") against such a file.
|
||||
--
|
||||
@@ -223,6 +221,28 @@ WHERE lower(name) IN ('rootblogname','noteblogname','type')
|
||||
HAVING COUNT(*) > 0;
|
||||
|
||||
|
||||
-- 1e. BLOGNAMES NOT RETIRED: a backup from between 2026-08-07 and 2026-09-28, when
|
||||
-- BlogNames still held the IDs and Blogs.BlogId was an
|
||||
-- unmaintained copy. Zero rows = good.
|
||||
--
|
||||
-- The current code resolves every Notes ID through Blogs.BlogId and never writes
|
||||
-- BlogNames, so against such a file new blogs get IDs that can collide with
|
||||
-- BlogNames' and every blog missing from Blogs.BlogId stays invisible to GetBlogs.
|
||||
--
|
||||
-- Fix: back up, then run retire-blognames.sql (after normalize-notes.sql if 1d
|
||||
-- also reported). It checks itself and changes nothing if a check fails.
|
||||
SELECT 'BlogNames still exists (' || type || ') -- run retire-blognames.sql' AS blognames_not_retired
|
||||
FROM sqlite_master
|
||||
WHERE lower(name) = 'blognames'
|
||||
UNION ALL
|
||||
SELECT 'Blogs.BlogId is not UNIQUE -- run retire-blognames.sql'
|
||||
WHERE NOT EXISTS (SELECT 1 FROM pragma_index_list('Blogs') WHERE name = 'ix_Blogs_BlogId' AND "unique" = 1)
|
||||
UNION ALL
|
||||
SELECT 'BlogId guard trigger missing: ' || t.name || ' -- run retire-blognames.sql'
|
||||
FROM (SELECT 'trg_Blogs_BlogId_NoDelete' AS name UNION ALL SELECT 'trg_Blogs_BlogId_Immutable') t
|
||||
WHERE NOT EXISTS (SELECT 1 FROM sqlite_master m WHERE m.type = 'trigger' AND m.name = t.name);
|
||||
|
||||
|
||||
-- ============================================================================
|
||||
-- SECTION 2 -- FIX (opt-in, additive only)
|
||||
--
|
||||
@@ -233,8 +253,8 @@ HAVING COUNT(*) > 0;
|
||||
-- subset. These are the 8 additive migration columns and nothing else; the
|
||||
-- likes high-water-mark reset is intentionally NOT included.
|
||||
--
|
||||
-- Nothing here addresses query 1d. The Notes integer schema is a data migration
|
||||
-- (normalize-notes.sql) and cannot be reached by adding columns.
|
||||
-- Nothing here addresses queries 1d or 1e. Those are data migrations
|
||||
-- (normalize-notes.sql, retire-blognames.sql) and cannot be reached by adding columns.
|
||||
-- ============================================================================
|
||||
|
||||
-- ALTER TABLE Posts ADD COLUMN PostType TEXT;
|
||||
|
||||
Reference in New Issue
Block a user