Compare commits
17
Commits
3d61cbb6ea
..
master
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
5af284a49f | ||
|
|
8fe2ffeb96 | ||
|
|
c8c43c4918 | ||
|
|
2948a4aff0 | ||
|
|
b8231d6a4c | ||
|
|
bbf05b3863 | ||
|
|
ea2afc9d35 | ||
|
|
7bf270e47d | ||
|
|
58d7b1d05e | ||
|
|
f41957fd2f | ||
|
|
a28c5cc9ec | ||
|
|
d6a96f7885 | ||
|
|
864b468d96 | ||
|
|
e6a3efba5b | ||
|
|
34da632e6a | ||
|
|
954ec353a5 | ||
|
|
c9530a3718 |
@@ -31,7 +31,7 @@ dotnet run -- --test [blogname] [postID] # Test API for specific post
|
|||||||
- `--test [blogname] [postID]`: Test API note collection
|
- `--test [blogname] [postID]`: Test API note collection
|
||||||
- `--posts`: Export post blogs to file
|
- `--posts`: Export post blogs to file
|
||||||
- `--blogs`: Export blog list to file
|
- `--blogs`: Export blog list to file
|
||||||
- `--collect [0|1] [datetime] [blogname]`: Collect notes for posts in DB. Optional `blogname` restricts the run to one blog (exact match), e.g. `--collect 1 zomb-eh`
|
- `--collect [0|1] [datetime] [blogname]`: Collect notes for posts in DB. Optional `blogname` restricts the run to one blog (exact match), e.g. `--collect 1 zomb-eh`. Add `--force` to ignore the periodic re-collect cooldown so already-collected posts are re-queued immediately (mode 1 only). Add `--fromDate <datetime>` / `--toDate <datetime>` to only re-queue already-collected posts whose original PostDate is on/after / on/before that date (mode 1 only; either or both may be given; applies with or without `--force`)
|
||||||
- `--blogsR`: Export reply blogs to file
|
- `--blogsR`: Export reply blogs to file
|
||||||
- `--blogsO [start] [stop]`: Export blogs within range
|
- `--blogsO [start] [stop]`: Export blogs within range
|
||||||
|
|
||||||
|
|||||||
@@ -49,28 +49,40 @@ say nothing about the item being fetched, so they must not be recorded as per-it
|
|||||||
|
|
||||||
### `Notes` Stores Integer IDs, Not Names
|
### `Notes` Stores Integer IDs, Not Names
|
||||||
As of 2026-08-07 `Notes.RootBlogName`, `NoteBlogName` and `Type` are gone, replaced by
|
As of 2026-08-07 `Notes.RootBlogName`, `NoteBlogName` and `Type` are gone, replaced by
|
||||||
`RootBlogId`, `NoteBlogId` and `TypeId` resolving through the `BlogNames` and `NoteTypes`
|
`RootBlogId`, `NoteBlogId` and `TypeId`. Blog IDs resolve through `Blogs.BlogId`, and
|
||||||
lookup tables. There is no compatibility view — naming an old column is a hard SQLite
|
types through the `NoteTypes` lookup table. There is no compatibility view: naming an old
|
||||||
error, so unlike `IsActive` this is a hard cut with no runtime probe. Full detail in
|
column is a hard SQLite error, so unlike `IsActive` this is a hard cut with no runtime
|
||||||
`URLNotesGrabberCORE/TL.db.md`.
|
probe. Full detail in `URLNotesGrabberCORE/TL.db.md`.
|
||||||
|
|
||||||
- **Joining `Notes` to `Blogs` goes through `Blogs.BlogId`**, not `BlogNames`:
|
- **`Blogs.BlogId` is the only blog-ID authority (since 2026-09-28).** IDs used to live in a
|
||||||
`FROM Blogs B INNER JOIN Notes N ON N.NoteBlogId = B.BlogId`. Routing it through
|
`BlogNames` table with an unmaintained copy in `Blogs.BlogId`. The copy drifted and hid
|
||||||
`BlogNames` adds a hop and ends in the text comparison the migration removed
|
12k blogs from `GetBlogs`, so `retire-blognames.sql` moved the authority into `Blogs`
|
||||||
- **Joining `Notes` to `Posts` is the opposite** — `Posts` has only `BlogName`, so it must
|
and **dropped `BlogNames` entirely**. There is no compatibility view, so naming it is
|
||||||
go through `BlogNames` (`GetRepliesWithFilledText`). This is the only such join
|
`no such table`. Do not recreate it
|
||||||
- **Resolve a name by filtering the lookup, never by scanning `Notes`**:
|
- **Joining `Notes` to `Blogs`**: `FROM Blogs B INNER JOIN Notes N ON N.NoteBlogId = B.BlogId`
|
||||||
`WHERE NoteBlogId = (SELECT BlogId FROM BlogNames WHERE BlogName = @name)`. The subquery
|
- **Joining `Notes` to `Posts` also goes through `Blogs`**, since `Posts` has only
|
||||||
is a unique-index probe on 20k rows and does not show against the 1.18M-row table
|
`BlogName`: `Posts P JOIN Blogs RB ON RB.BlogName = P.BlogName JOIN Notes N ON
|
||||||
- **`AddNote` registers both blog names *and* the note type** with `INSERT OR IGNORE`
|
N.RootBlogId = RB.BlogId` (`GetRepliesWithFilledText`)
|
||||||
before inserting, all in one transaction. `NoteTypes` is a table rather than a `CHECK`
|
- **Resolve a name by filtering `Blogs`, never by scanning `Notes`**:
|
||||||
constraint precisely so an unseen type is an `INSERT`; without that registration it
|
`WHERE NoteBlogId = (SELECT BlogId FROM Blogs WHERE BlogName = @name)`. The subquery is a
|
||||||
would resolve to `NULL` and fail the `NOT NULL` on `TypeId`, losing the note
|
primary-key probe and does not show against the 1.2M-row table
|
||||||
- **`Blogs.BlogId` is NULL on 168,202 of 188,620 rows** — every blog that has never
|
- **`AddNote` registers both blogs *and* the note type** before inserting, all in one
|
||||||
appeared in a note. An inner join on it silently drops them. Correct for engagement
|
transaction. `RegisterBlog` does `INSERT OR IGNORE` into `Blogs`, then assigns
|
||||||
queries, wrong for anything listing the registry
|
`BlogId = MAX(BlogId) + 1` where it is NULL. Unlike `AddBlog`, it does not skip `deact`
|
||||||
- **IDs are stable and must never be renumbered.** They are stored in 1.18M `Notes` rows.
|
names, because a note by a deactivated blog still needs an ID. `NoteTypes` is a table
|
||||||
A blog renamed upstream gets a new `BlogNames` row, not an edited one
|
rather than a `CHECK` constraint precisely so an unseen type is an `INSERT`. Without
|
||||||
|
that registration a type would resolve to `NULL` and fail the `NOT NULL` on `TypeId`,
|
||||||
|
losing the note
|
||||||
|
- **Assigning a `BlogId` is bookkeeping and must not move `DateModified`**
|
||||||
|
- **`Blogs.BlogId` is NULL on ~166k of ~199k rows**, every blog that has never appeared in
|
||||||
|
a note. An inner join on it silently drops them. Correct for engagement queries, wrong
|
||||||
|
for anything listing the registry. `ix_Blogs_BlogId` is `UNIQUE`, which allows many NULLs
|
||||||
|
- **IDs are stable and must never be renumbered.** They are stored in 1.2M `Notes` rows.
|
||||||
|
Triggers `trg_Blogs_BlogId_NoDelete` and `trg_Blogs_BlogId_Immutable` abort any
|
||||||
|
`DELETE` of a `Blogs` row that has a `BlogId`, and any change to its `BlogId` or
|
||||||
|
`BlogName`. A blog renamed upstream gets a new row. Remove a blog with `IsActive = 0`.
|
||||||
|
These triggers are also what make `MAX(BlogId) + 1` safe: no ID can ever be freed for
|
||||||
|
reuse
|
||||||
- Prefer `TypeId = (SELECT TypeId FROM NoteTypes WHERE Type = 'reply')` over a hardcoded
|
- Prefer `TypeId = (SELECT TypeId FROM NoteTypes WHERE Type = 'reply')` over a hardcoded
|
||||||
ID. A negated `TypeId NOT IN (SELECT …)` is only correct because `TypeId` is `NOT NULL`
|
ID. A negated `TypeId NOT IN (SELECT …)` is only correct because `TypeId` is `NOT NULL`
|
||||||
- Duplicate-key detection uses `IsNotesDuplicateKey`, which matches the constraint and the
|
- Duplicate-key detection uses `IsNotesDuplicateKey`, which matches the constraint and the
|
||||||
|
|||||||
+385
-70
@@ -1,71 +1,90 @@
|
|||||||
<?xml version="1.0" encoding="UTF-8"?><sqlb_project><db path="D:/NextCloud/C#/URLNotesGrabberCORE/URLNotesGrabberCORE/TL.db" readonly="0" foreign_keys="1" case_sensitive_like="0" temp_store="0" wal_autocheckpoint="1000" synchronous="2"/><attached/><window><main_tabs open="structure browser pragmas query" current="3"/></window><tab_structure><column_width id="0" width="300"/><column_width id="1" width="0"/><column_width id="2" width="100"/><column_width id="3" width="4486"/><column_width id="4" width="0"/><expanded_item id="0" parent="1"/><expanded_item id="1" parent="1"/><expanded_item id="2" parent="1"/><expanded_item id="3" parent="1"/></tab_structure><tab_browse><table title="Posts" custom_title="0" dock_id="4" table="4,5:mainPosts"/><dock_state state="000000ff00000000fd0000000100000002000005470000029afc0100000006fb000000160064006f0063006b00420072006f00770073006500310100000000000004a10000000000000000fb000000160064006f0063006b00420072006f00770073006500320100000000000004a10000000000000000fb000000160064006f0063006b00420072006f00770073006500330100000000000004a10000000000000000fb000000160064006f0063006b00420072006f00770073006500350100000000000005f40000000000000000fb000000160064006f0063006b00420072006f00770073006500340100000000000005470000011100fffffffb000000160064006f0063006b00420072006f00770073006500340100000000000005f40000000000000000000005470000000000000004000000040000000800000008fc00000000"/><default_encoding codec=""/><browse_table_settings><table schema="main" name="ApiKeyPoolMeta" show_row_id="0" encoding="" plot_x_axis="" unlock_view_pk="_rowid_" freeze_columns="0"><sort/><column_widths><column index="1" value="29"/><column index="2" value="64"/></column_widths><filter_values/><conditional_formats/><row_id_formats/><display_formats/><hidden_columns/><plot_y_axes/><global_filter/></table><table schema="main" name="Blogs" show_row_id="0" encoding="" plot_x_axis="" unlock_view_pk="_rowid_" freeze_columns="0"><sort><column index="7" mode="1"/></sort><column_widths><column index="1" value="257"/><column index="2" value="108"/><column index="3" value="63"/><column index="4" value="156"/><column index="5" value="60"/><column index="6" value="81"/><column index="7" value="85"/><column index="8" value="156"/><column index="9" value="156"/><column index="10" value="151"/><column index="11" value="125"/><column index="12" value="129"/></column_widths><filter_values><column index="4" value="1"/><column index="7" value=">2026-05-27 17:22:36"/></filter_values><conditional_formats/><row_id_formats/><display_formats/><hidden_columns/><plot_y_axes/><global_filter/></table><table schema="main" name="Notes" show_row_id="0" encoding="" plot_x_axis="" unlock_view_pk="_rowid_" freeze_columns="0"><sort/><column_widths><column index="1" value="198"/><column index="2" value="144"/><column index="3" value="251"/><column index="4" value="84"/><column index="5" value="53"/><column index="6" value="300"/><column index="7" value="116"/><column index="8" value="152"/><column index="9" value="89"/><column index="10" value="63"/></column_widths><filter_values/><conditional_formats/><row_id_formats/><display_formats/><hidden_columns/><plot_y_axes/><global_filter/></table><table schema="main" name="Posts" show_row_id="0" encoding="" plot_x_axis="" unlock_view_pk="_rowid_" freeze_columns="0"><sort><column index="14" mode="1"/></sort><column_widths><column index="1" value="236"/><column index="2" value="144"/><column index="3" value="126"/><column index="4" value="32"/><column index="5" value="32"/><column index="6" value="32"/><column index="7" value="32"/><column index="8" value="32"/><column index="9" value="0"/><column index="10" value="0"/><column index="11" value="0"/><column index="12" value="243"/><column index="13" value="300"/><column index="14" value="53"/><column index="15" value="37351"/><column index="16" value="300"/><column index="17" value="41"/><column index="18" value="75"/><column index="19" value="96"/><column index="20" value="300"/><column index="21" value="96"/><column index="22" value="300"/><column index="23" value="300"/><column index="24" value="548"/><column index="25" value="60"/><column index="26" value="213"/><column index="27" value="532"/><column index="28" value="152"/><column index="29" value="152"/><column index="30" value="69"/><column index="31" value="63"/></column_widths><filter_values><column index="2" value="0"/><column index="1" value="734568371821084672"/></filter_values><conditional_formats/><row_id_formats/><display_formats/><hidden_columns><column index="9" value="1"/><column index="10" value="1"/><column index="11" value="1"/></hidden_columns><plot_y_axes/><global_filter/></table></browse_table_settings></tab_browse><tab_sql><sql name="SQL 1">UPDATE Posts
|
<?xml version="1.0" encoding="UTF-8"?><sqlb_project><db path="C:/Users/jim/Nextcloud/C#/URLNotesGrabberCORE/URLNotesGrabberCORE/TL.db" readonly="0" foreign_keys="1" case_sensitive_like="0" temp_store="0" wal_autocheckpoint="1000" synchronous="2"/><attached/><window><main_tabs open="structure browser pragmas query" current="3"/></window><tab_structure><column_width id="0" width="300"/><column_width id="1" width="0"/><column_width id="2" width="100"/><column_width id="3" width="4486"/><column_width id="4" width="0"/><expanded_item id="0" parent="1"/><expanded_item id="1" parent="1"/><expanded_item id="2" parent="1"/><expanded_item id="3" parent="1"/></tab_structure><tab_browse><table title="Notes" custom_title="0" dock_id="4" table="4,5:mainNotes"/><dock_state state="000000ff00000000fd0000000100000002000005470000029efc0100000006fb000000160064006f0063006b00420072006f00770073006500310100000000000004a10000000000000000fb000000160064006f0063006b00420072006f00770073006500320100000000000004a10000000000000000fb000000160064006f0063006b00420072006f00770073006500330100000000000004a10000000000000000fb000000160064006f0063006b00420072006f00770073006500350100000000000005f40000000000000000fb000000160064006f0063006b00420072006f00770073006500340100000000000005470000011100fffffffb000000160064006f0063006b00420072006f00770073006500340100000000000005f40000000000000000000005470000000000000004000000040000000800000008fc00000000"/><default_encoding codec=""/><browse_table_settings><table schema="main" name="ApiKeyPoolMeta" show_row_id="0" encoding="" plot_x_axis="" unlock_view_pk="_rowid_" freeze_columns="0"><sort/><column_widths><column index="1" value="29"/><column index="2" value="64"/></column_widths><filter_values/><conditional_formats/><row_id_formats/><display_formats/><hidden_columns/><plot_y_axes/><global_filter/></table><table schema="main" name="Blogs" show_row_id="0" encoding="" plot_x_axis="" unlock_view_pk="_rowid_" freeze_columns="0"><sort><column index="7" mode="1"/></sort><column_widths><column index="1" value="257"/><column index="2" value="108"/><column index="3" value="63"/><column index="4" value="156"/><column index="5" value="60"/><column index="6" value="81"/><column index="7" value="85"/><column index="8" value="156"/><column index="9" value="156"/><column index="10" value="151"/><column index="11" value="125"/><column index="12" value="129"/></column_widths><filter_values><column index="4" value="1"/><column index="7" value=">2026-05-27 17:22:36"/></filter_values><conditional_formats/><row_id_formats/><display_formats/><hidden_columns/><plot_y_axes/><global_filter/></table><table schema="main" name="Notes" show_row_id="0" encoding="" plot_x_axis="" unlock_view_pk="_rowid_" freeze_columns="0"><sort><column index="4" mode="1"/></sort><column_widths><column index="1" value="81"/><column index="2" value="148"/><column index="3" value="83"/><column index="4" value="85"/><column index="5" value="56"/><column index="6" value="300"/><column index="7" value="156"/><column index="8" value="156"/><column index="9" value="156"/><column index="10" value="63"/></column_widths><filter_values><column index="2" value="4370"/></filter_values><conditional_formats/><row_id_formats/><display_formats/><hidden_columns/><plot_y_axes/><global_filter/></table><table schema="main" name="Posts" show_row_id="0" encoding="" plot_x_axis="" unlock_view_pk="_rowid_" freeze_columns="0"><sort><column index="14" mode="1"/></sort><column_widths><column index="1" value="241"/><column index="2" value="148"/><column index="3" value="126"/><column index="4" value="300"/><column index="5" value="75"/><column index="6" value="187"/><column index="7" value="159"/><column index="8" value="75"/><column index="9" value="0"/><column index="10" value="0"/><column index="11" value="0"/><column index="12" value="249"/><column index="13" value="300"/><column index="14" value="53"/><column index="15" value="300"/><column index="16" value="300"/><column index="17" value="41"/><column index="18" value="75"/><column index="19" value="96"/><column index="20" value="300"/><column index="21" value="96"/><column index="22" value="300"/><column index="23" value="300"/><column index="24" value="300"/><column index="25" value="60"/><column index="26" value="218"/><column index="27" value="300"/><column index="28" value="156"/><column index="29" value="156"/><column index="30" value="69"/><column index="31" value="63"/></column_widths><filter_values><column index="1" value="137735301451"/></filter_values><conditional_formats/><row_id_formats/><display_formats/><hidden_columns><column index="9" value="1"/><column index="10" value="1"/><column index="11" value="1"/></hidden_columns><plot_y_axes/><global_filter/></table></browse_table_settings></tab_browse><tab_sql><sql name="Mark Blogs">select *
|
||||||
SET HasNotesGathered = 0
|
|
||||||
WHERE (BlogName, PostID) IN (
|
|
||||||
SELECT p.BlogName, p.PostID
|
|
||||||
FROM Posts p
|
|
||||||
WHERE p.HasNotesGathered = 1
|
|
||||||
AND P.notesGatheredDatetime < 1774294520
|
|
||||||
AND EXISTS (
|
|
||||||
SELECT 1
|
|
||||||
FROM Notes n
|
|
||||||
WHERE n.PostID = p.PostID
|
|
||||||
AND n.RootBlogName = p.BlogName
|
|
||||||
--AND n.Type NOT IN ('reblog', 'reply')
|
|
||||||
)
|
|
||||||
ORDER BY P.PostDate ASC
|
|
||||||
--LIMIT 500
|
|
||||||
);</sql><sql name="Mark Blogs">select *
|
|
||||||
from Blogs
|
from Blogs
|
||||||
--update blogs set HasBeenOutput = 1
|
--update blogs set HasBeenOutput = 1
|
||||||
where HasBeenOutput = 0
|
where HasBeenOutput = 0
|
||||||
AND
|
AND
|
||||||
blogname in
|
blogname in
|
||||||
(
|
('udontn33dh1m',
|
||||||
'teaberrybee',
|
'tyrantsxblood',
|
||||||
'reddevilgoddesstoo',
|
'sentry-34',
|
||||||
'waywardog13',
|
'deathcabforfrankie',
|
||||||
'wzjustbrowsing-blog',
|
'abheith-sasta',
|
||||||
'lewerta',
|
'kuwaiikittenghost',
|
||||||
'nudenymph',
|
'kansasmud',
|
||||||
'caylachief'
|
'03diesel',
|
||||||
|
'itzameallieee',
|
||||||
)</sql><sql name="New Notes">select P.slug, N.replyText, n.RootBlogName, n.PostID, NoteBlogName || '.tumblr.com' as NoteBlogName, DatetimeCrawled, TimeStamp, type, n.RootBlogName || '.tumblr.com/post/' || n.postid, datetime(timestamp, 'unixepoch')
|
'fireball-temptations',
|
||||||
from Notes N inner join Posts P on p.PostID = n.PostID
|
'mamaisamess',
|
||||||
|
'906raised-and-dogobsessed',
|
||||||
|
'the-queerist-wolf',
|
||||||
|
'counting-corpsess',
|
||||||
|
'aqueenbby',
|
||||||
|
'maybememoriesx',
|
||||||
|
'queenofnevers',
|
||||||
|
'obsidian-psyche',
|
||||||
|
'lilmissellexo',
|
||||||
|
'alittlebunny95',
|
||||||
|
'rage--and--grace',
|
||||||
|
'savage-deniz',
|
||||||
|
'daddyspuddleprincess',
|
||||||
|
'littledefenstration',
|
||||||
|
'bearded-snorlax',
|
||||||
|
'thosesummerskiess',
|
||||||
|
'tubadtoph',
|
||||||
|
'lieutenant-dan-ice-cream',
|
||||||
|
'brittvnybitch',
|
||||||
|
'a-smol-gayologist',
|
||||||
|
'sum1random',
|
||||||
|
'samsternelly',
|
||||||
|
'littlemouseylauren',
|
||||||
|
'princessleiaorgasma',
|
||||||
|
'bloodstaineddkisses',
|
||||||
|
'letsfacerealitybabe',
|
||||||
|
'x--marks--thespot',
|
||||||
|
'space-and-suffering',
|
||||||
|
'rinarootski',
|
||||||
|
'thiccandtired',
|
||||||
|
'fvcking-scvmbag',
|
||||||
|
'fullblownwizard',
|
||||||
|
'bigjewface',
|
||||||
|
'unleash-the-krayken',
|
||||||
|
'bumpintheroad',
|
||||||
|
'liltexasjedii',
|
||||||
|
'nawtydude',
|
||||||
|
'queenpeachqueen',
|
||||||
|
'the-clansman',
|
||||||
|
'balmain-bxtch'
|
||||||
|
)</sql><sql name="New Notes">select P.slug, N.replyText, rbn.BlogName as RootBlogName, n.PostID, nbn.BlogName || '.tumblr.com' as NoteBlogName, DatetimeCrawled, TimeStamp, nt.Type, rbn.BlogName || '.tumblr.com/post/' || n.postid, datetime(timestamp, 'unixepoch')
|
||||||
|
from Notes N
|
||||||
|
inner join Posts P on p.PostID = n.PostID
|
||||||
|
inner join Blogs rbn on rbn.BlogId = n.RootBlogId
|
||||||
|
inner join Blogs nbn on nbn.BlogId = n.NoteBlogId
|
||||||
|
inner join NoteTypes nt on nt.TypeId = n.TypeId
|
||||||
where
|
where
|
||||||
DatetimeCrawled > '2026-08-07 11:47:22' and type like 'r%'
|
DatetimeCrawled > '2026-08-07 11:47:22' and nt.Type like 'r%'
|
||||||
and P.IsActive = 1
|
and P.IsActive = 1
|
||||||
order by n.DatetimeCrawled</sql><sql name="Pull Blogs">SELECT distinct
|
order by n.DatetimeCrawled</sql><sql name="SQL 7">WITH ReplyCounts AS (
|
||||||
'''' || blogname || ''',',
|
|
||||||
blogs.*
|
|
||||||
, blogname || '.tumblr.com'
|
|
||||||
FROM
|
|
||||||
Blogs
|
|
||||||
inner JOIN
|
|
||||||
Notes on notes.noteBlogName = blogs.BlogName
|
|
||||||
WHERE
|
|
||||||
HasBeenOutput = 0 and type = 'reblog'
|
|
||||||
order by
|
|
||||||
Notes.Type desc,
|
|
||||||
DateAdded desc
|
|
||||||
LIMIT 100;</sql><sql name="SQL 7">WITH ReplyCounts AS (
|
|
||||||
SELECT
|
SELECT
|
||||||
NoteBlogName,
|
NoteBlogId,
|
||||||
COUNT(DISTINCT replyText) AS DistinctReplyCount
|
COUNT(DISTINCT replyText) AS DistinctReplyCount
|
||||||
FROM Notes
|
FROM Notes
|
||||||
where replyText <> '.'
|
where replyText <> '.'
|
||||||
GROUP BY NoteBlogName
|
GROUP BY NoteBlogId
|
||||||
)
|
)
|
||||||
SELECT
|
SELECT
|
||||||
n.RootBlogName || '.tumblr.com/post/' || n.PostID AS PostURL, postid,
|
rbn.BlogName || '.tumblr.com/post/' || n.PostID AS PostURL, postid,
|
||||||
n.NoteBlogName,
|
nbn.BlogName AS NoteBlogName,
|
||||||
n.replyText,
|
n.replyText,
|
||||||
c.DistinctReplyCount
|
c.DistinctReplyCount
|
||||||
FROM Notes n
|
FROM Notes n
|
||||||
JOIN ReplyCounts c ON n.NoteBlogName = c.NoteBlogName
|
JOIN ReplyCounts c ON n.NoteBlogId = c.NoteBlogId
|
||||||
where replyText <> '.' and type <> 'reply'
|
JOIN Blogs rbn ON rbn.BlogId = n.RootBlogId
|
||||||
--AND N.NoteBlogName NOT IN ( 'roadblocker21', 'thesaddemon666', 'edwardabbeyhoffman', 'tattedsoldier20', 'zomb-eh', 'animalistic13', 'indken', 'maccloud1592',
|
JOIN Blogs nbn ON nbn.BlogId = n.NoteBlogId
|
||||||
|
JOIN NoteTypes t ON t.TypeId = n.TypeId
|
||||||
|
where replyText <> '.' and t.Type <> 'reply'
|
||||||
|
--AND nbn.BlogName NOT IN ( 'roadblocker21', 'thesaddemon666', 'edwardabbeyhoffman', 'tattedsoldier20', 'zomb-eh', 'animalistic13', 'indken', 'maccloud1592',
|
||||||
--'moss-wizard', 'supertrucker12682', 'exploringthrupics', 'padeyepete' )
|
--'moss-wizard', 'supertrucker12682', 'exploringthrupics', 'padeyepete' )
|
||||||
order by c.DistinctReplyCount desc, n.NoteBlogName, n.DateModified desc, replyText, RootBlogName, PostID</sql><sql name="Collect">WITH PostsWithCount AS ( SELECT P.BlogName, P.PostID, 1925013599 AS LatestNoteTimestamp, P.NotesGatheredDateTime, COUNT(P.PostID) OVER(PARTITION BY P.BlogName) AS CNT, P.HasNotesGathered, P.NotFound, P.PostDate FROM Posts P WHERE COALESCE(P.IsActive, 1) = 1 ), Unioned AS ( SELECT BlogName, PostID, LatestNoteTimestamp, NotesGatheredDateTime, CNT, PostDate FROM PostsWithCount WHERE NotFound = 0 AND HasNotesGathered = 0 UNION SELECT BlogName, PostID, LatestNoteTimestamp, NotesGatheredDateTime, CNT, PostDate FROM PostsWithCount WHERE BlogName = 'zomb-eh' AND NotFound = 0 AND NotesGatheredDateTime < unixepoch('now', 'localtime', '-3 days') ) SELECT U.BlogName, U.PostID, U.LatestNoteTimestamp, U.NotesGatheredDateTime, U.CNT FROM Unioned U WHERE (U.NotesGatheredDateTime < 1786134037 OR U.NotesGatheredDateTime IS NULL) ORDER BY U.NotesGatheredDateTime, U.PostDate DESC, U.BlogName, U.PostID;</sql><sql name="Del Posts">delete from posts where postid in
|
order by c.DistinctReplyCount desc, nbn.BlogName, n.DateModified desc, replyText, rbn.BlogName, PostID</sql><sql name="Collect">WITH PostsWithCount AS ( SELECT P.BlogName, P.PostID, 1925013599 AS LatestNoteTimestamp, P.NotesGatheredDateTime, COUNT(P.PostID) OVER(PARTITION BY P.BlogName) AS CNT, P.HasNotesGathered, P.NotFound, P.PostDate FROM Posts P WHERE COALESCE(P.IsActive, 1) = 1 ), Unioned AS ( SELECT BlogName, PostID, LatestNoteTimestamp, NotesGatheredDateTime, CNT, PostDate FROM PostsWithCount WHERE NotFound = 0 AND HasNotesGathered = 0 UNION SELECT BlogName, PostID, LatestNoteTimestamp, NotesGatheredDateTime, CNT, PostDate FROM PostsWithCount WHERE BlogName = 'zomb-eh' AND NotFound = 0 AND NotesGatheredDateTime < unixepoch('now', 'localtime', '-3 days') ) SELECT U.BlogName, U.PostID, U.LatestNoteTimestamp, U.NotesGatheredDateTime, U.CNT FROM Unioned U WHERE (U.NotesGatheredDateTime < 1787237598 OR U.NotesGatheredDateTime IS NULL) ORDER BY U.NotesGatheredDateTime, U.PostDate DESC, U.BlogName, U.PostID;</sql><sql name="Del Posts">delete from posts where postid in
|
||||||
(
|
(
|
||||||
'741662499571728384',
|
'741662499571728384',
|
||||||
178892849664,
|
178892849664,
|
||||||
@@ -73,28 +92,324 @@ order by c.DistinctReplyCount desc, n.NoteBlogName, n.DateModified desc, replyTe
|
|||||||
177012868749,
|
177012868749,
|
||||||
169950081964,
|
169950081964,
|
||||||
755440787056099328
|
755440787056099328
|
||||||
)</sql><sql name="notes NO post*">select *
|
)</sql><sql name="notes NO post">select *
|
||||||
-- delete
|
-- delete
|
||||||
from notes
|
from notes
|
||||||
where postid not in (select distinct postid from posts where IsActive = 1)</sql><sql name="SQL 9">SELECT
|
where postid not in (select distinct postid from posts where IsActive = 1)</sql><sql name="Pull Blogs">-- ============================================================================
|
||||||
*
|
-- blogs-added-after-august-2026-with-reblog-or-reply.sql
|
||||||
FROM
|
--
|
||||||
POSTS P
|
-- Purpose: Of all blogs in Blogs, find the ones that show up in Notes as the
|
||||||
WHERE
|
-- engager (NoteBlogId) on a 'reblog' or 'reply' note, sorted by
|
||||||
P.ByLikes = 1
|
-- DateAdded. (Originally scoped to "added after August 2026" --
|
||||||
AND
|
-- that cutoff is now removed per request; QUERY 2 shows how to put
|
||||||
P.DateCreated > '2026-05-26 17:47:32'
|
-- a date floor back if needed.)
|
||||||
ORDER BY
|
--
|
||||||
P.DateCreated desc</sql><sql name="SQL 13">update posts set IsActive = 0 where blogname IN ( 'shoebiedoo', 'redheaded-girlygirl', 'xlittle-ghost' )</sql><sql name="SQL 14*">update Posts␍
|
-- Read-only. No INSERT/UPDATE/DELETE/DDL anywhere in this file.
|
||||||
set IsActive = 0␍
|
--
|
||||||
where postid in␍
|
-- How to use (DB Browser for SQLite):
|
||||||
(␍
|
-- 1. File > Open Database -> TL.db
|
||||||
␍
|
-- 2. Execute SQL tab. Each QUERY below is independent; Ctrl+Enter runs just
|
||||||
␍
|
-- the one your cursor is in.
|
||||||
'731937314675310592'␍
|
--
|
||||||
␍
|
-- The join, once:
|
||||||
|
-- "Added after August 2026" filters Blogs.DateAdded. "Has a reblog/reply
|
||||||
|
-- note" means the blog is the engager, which is NoteBlogId -- not
|
||||||
|
-- RootBlogId, which is the blog that *owns* the post being reacted to
|
||||||
|
-- (see find-notes-on-inactive-posts.sql for that side). Per TL.db.md, the
|
||||||
|
-- Blogs<->Notes join is a single integer hop and should not be routed
|
||||||
|
-- through Blogs:
|
||||||
|
-- Blogs.BlogId = Notes.NoteBlogId
|
||||||
|
-- EXISTS is used rather than a JOIN so a blog with many qualifying notes
|
||||||
|
-- still contributes one output row.
|
||||||
|
--
|
||||||
|
-- Excluding notes on an inactive post: same shape as
|
||||||
|
-- find-notes-on-inactive-posts.sql -- Notes only carries RootBlogId (an
|
||||||
|
-- integer), so reaching Posts.IsActive needs the one text hop the rest of
|
||||||
|
-- this file avoids: RootBlogId -> Blogs.BlogName = Posts.BlogName,
|
||||||
|
-- matched on PostID. It's a LEFT JOIN, not an inner one: only 3,867 of
|
||||||
|
-- 20,430 root/engager blogs have any stored Posts rows at all (TL.db.md),
|
||||||
|
-- so most reblog/reply notes have no Posts row to check and must be kept,
|
||||||
|
-- not dropped by an inner join. COALESCE(p.IsActive, 1) = 1 keeps a note
|
||||||
|
-- unless its post is explicitly IsActive = 0 -- NULL (no Posts row, or a
|
||||||
|
-- stored row with no flag written) means live, per the schema's own
|
||||||
|
-- convention (TL.db.md, "Posts.IsActive and Notes.IsActive"). This is a
|
||||||
|
-- big filter in practice: of the blogs that qualified before it, most
|
||||||
|
-- have every one of their reblog/reply notes pointing at a since-removed
|
||||||
|
-- post, not just some -- verified against the live data, not assumed.
|
||||||
|
--
|
||||||
|
-- On DateAdded: this column is not written consistently -- most rows hold
|
||||||
|
-- ISO 'yyyy-MM-dd HH:mm:ss', but 17k+ hold US 'M/d/yy' from a 2025 bulk
|
||||||
|
-- import (see TL.db.md, "DateAdded is not written consistently"). As text,
|
||||||
|
-- those two shapes do not sort or compare against each other correctly, so
|
||||||
|
-- QUERY 0 normalises both to an ISO date before filtering. In the live data
|
||||||
|
-- every US-format row predates August 2026 anyway (only '12/23/25' and
|
||||||
|
-- '12/24/25' occur), so this makes no difference to the current answer --
|
||||||
|
-- it's here so the query stays correct if that ever changes.
|
||||||
|
-- ============================================================================
|
||||||
|
|
||||||
|
|
||||||
|
-- ----------------------------------------------------------------------------
|
||||||
|
-- QUERY 0 / THE ANSWER -- one row per qualifying blog, sorted by DateAdded
|
||||||
|
-- descending (normalised -- see the note above). No date cutoff, but now
|
||||||
|
-- scoped to HasBeenOutput = 0 AND IsActive = 1. 4,739 rows in the live
|
||||||
|
-- data.
|
||||||
|
--
|
||||||
|
-- earliest_reblog_or_reply_utc is the MIN(TimeStamp) among this blog's
|
||||||
|
-- reblog-or-reply notes (either type counts -- see the column name).
|
||||||
|
-- Getting this meant switching QUERY 0 from EXISTS to an inner JOIN +
|
||||||
|
-- GROUP BY: EXISTS can only tell you a qualifying row is present, not
|
||||||
|
-- aggregate over which ones. No CASE is needed inside the MIN() because
|
||||||
|
-- the WHERE below already restricts the joined rows to reblog/reply, so
|
||||||
|
-- every row a blog brings into the aggregate is one this column should
|
||||||
|
-- consider. A blog appears exactly once, same as before, and this column
|
||||||
|
-- is never NULL for a row that's in the result at all (an earlier
|
||||||
|
-- revision aggregated reblog only, which left it NULL for the 181 blogs
|
||||||
|
-- that had replies but no reblogs).
|
||||||
|
-- ----------------------------------------------------------------------------
|
||||||
|
WITH BlogsSplit AS (
|
||||||
|
SELECT
|
||||||
|
b.BlogId,
|
||||||
|
b.BlogName,
|
||||||
|
b.DateAdded,
|
||||||
|
CASE WHEN b.DateAdded LIKE '____-__-__%' THEN 1 ELSE 0 END AS IsIso,
|
||||||
|
-- for the US 'M/d/yy' shape only: everything after the first '/'
|
||||||
|
substr(b.DateAdded, instr(b.DateAdded, '/') + 1) AS RestAfterMonth
|
||||||
|
FROM Blogs b
|
||||||
|
WHERE b.BlogId IS NOT NULL and HasBeenOutput = 0 and IsActive = 1 -- a blog can only match Notes if it has one
|
||||||
|
),
|
||||||
|
BlogsNorm AS (
|
||||||
|
SELECT
|
||||||
|
BlogId,
|
||||||
|
BlogName,
|
||||||
|
DateAdded,
|
||||||
|
CASE
|
||||||
|
WHEN IsIso = 1 THEN date(DateAdded)
|
||||||
|
ELSE date(
|
||||||
|
'20' || substr(RestAfterMonth, instr(RestAfterMonth, '/') + 1) || '-' ||
|
||||||
|
substr('00' || substr(DateAdded, 1, instr(DateAdded, '/') - 1), -2) || '-' ||
|
||||||
|
substr('00' || substr(RestAfterMonth, 1, instr(RestAfterMonth, '/') - 1), -2)
|
||||||
|
)
|
||||||
|
END AS DateAddedNorm
|
||||||
|
FROM BlogsSplit
|
||||||
)
|
)
|
||||||
|
SELECT
|
||||||
|
bn.BlogId,
|
||||||
|
bn.BlogName,
|
||||||
|
bn.DateAdded,
|
||||||
|
bn.DateAddedNorm,
|
||||||
|
datetime(MIN(n.TimeStamp), 'unixepoch') AS earliest_reblog_or_reply_utc
|
||||||
|
FROM BlogsNorm bn
|
||||||
|
JOIN Notes n ON n.NoteBlogId = bn.BlogId
|
||||||
|
JOIN NoteTypes t ON t.TypeId = n.TypeId
|
||||||
|
JOIN Blogs root_bn ON root_bn.BlogId = n.RootBlogId
|
||||||
|
LEFT JOIN Posts p ON p.BlogName = root_bn.BlogName
|
||||||
|
AND p.PostID = n.PostID
|
||||||
|
WHERE t.Type IN ('reblog')--, 'reply')
|
||||||
|
AND COALESCE(p.IsActive, 1) = 1 -- exclude notes on a post explicitly marked removed
|
||||||
|
GROUP BY bn.BlogId, bn.BlogName, bn.DateAdded, bn.DateAddedNorm
|
||||||
|
ORDER BY bn.DateAddedNorm desc;
|
||||||
|
|
||||||
</sql><current_tab id="7"/></tab_sql></sqlb_project>
|
|
||||||
|
-- ----------------------------------------------------------------------------
|
||||||
|
-- QUERY 1 -- same answer, with a per-blog breakdown of which type(s) fired
|
||||||
|
-- and how many. Useful once QUERY 0 has rows; redundant while it's empty.
|
||||||
|
-- ----------------------------------------------------------------------------
|
||||||
|
-- WITH BlogsSplit AS ( ... ), BlogsNorm AS ( ... ) -- reuse the CTEs above
|
||||||
|
--
|
||||||
|
-- SELECT
|
||||||
|
-- bn.BlogId,
|
||||||
|
-- bn.BlogName,
|
||||||
|
-- bn.DateAddedNorm,
|
||||||
|
-- SUM(CASE WHEN t.Type = 'reblog' THEN 1 ELSE 0 END) AS reblog_count,
|
||||||
|
-- SUM(CASE WHEN t.Type = 'reply' THEN 1 ELSE 0 END) AS reply_count
|
||||||
|
-- FROM BlogsNorm bn
|
||||||
|
-- JOIN Notes n ON n.NoteBlogId = bn.BlogId
|
||||||
|
-- JOIN NoteTypes t ON t.TypeId = n.TypeId
|
||||||
|
-- WHERE t.Type IN ('reblog', 'reply')
|
||||||
|
-- GROUP BY bn.BlogId, bn.BlogName, bn.DateAddedNorm
|
||||||
|
-- ORDER BY bn.DateAddedNorm;
|
||||||
|
|
||||||
|
|
||||||
|
-- ----------------------------------------------------------------------------
|
||||||
|
-- QUERY 2 -- put a date floor back, if wanted later.
|
||||||
|
-- Same as QUERY 0, with one extra line in the outer WHERE:
|
||||||
|
-- AND bn.DateAddedNorm >= '2026-09-01' -- or whatever cutoff
|
||||||
|
-- ----------------------------------------------------------------------------
|
||||||
|
-- ============================================================================
|
||||||
|
-- blogs-added-after-august-2026-with-reblog-or-reply.sql
|
||||||
|
--
|
||||||
|
-- Purpose: Of all blogs in Blogs, find the ones that show up in Notes as the
|
||||||
|
-- engager (NoteBlogId) on a 'reblog' or 'reply' note, sorted by
|
||||||
|
-- DateAdded. (Originally scoped to "added after August 2026" --
|
||||||
|
-- that cutoff is now removed per request; QUERY 2 shows how to put
|
||||||
|
-- a date floor back if needed.)
|
||||||
|
--
|
||||||
|
-- Read-only. No INSERT/UPDATE/DELETE/DDL anywhere in this file.
|
||||||
|
--
|
||||||
|
-- How to use (DB Browser for SQLite):
|
||||||
|
-- 1. File > Open Database -> TL.db
|
||||||
|
-- 2. Execute SQL tab. Each QUERY below is independent; Ctrl+Enter runs just
|
||||||
|
-- the one your cursor is in.
|
||||||
|
--
|
||||||
|
-- The join, once:
|
||||||
|
-- "Added after August 2026" filters Blogs.DateAdded. "Has a reblog/reply
|
||||||
|
-- note" means the blog is the engager, which is NoteBlogId -- not
|
||||||
|
-- RootBlogId, which is the blog that *owns* the post being reacted to
|
||||||
|
-- (see find-notes-on-inactive-posts.sql for that side). Per TL.db.md, the
|
||||||
|
-- Blogs<->Notes join is a single integer hop and should not be routed
|
||||||
|
-- through Blogs:
|
||||||
|
-- Blogs.BlogId = Notes.NoteBlogId
|
||||||
|
-- EXISTS is used rather than a JOIN so a blog with many qualifying notes
|
||||||
|
-- still contributes one output row.
|
||||||
|
--
|
||||||
|
-- Excluding notes on an inactive post: same shape as
|
||||||
|
-- find-notes-on-inactive-posts.sql -- Notes only carries RootBlogId (an
|
||||||
|
-- integer), so reaching Posts.IsActive needs the one text hop the rest of
|
||||||
|
-- this file avoids: RootBlogId -> Blogs.BlogName = Posts.BlogName,
|
||||||
|
-- matched on PostID. It's a LEFT JOIN, not an inner one: only 3,867 of
|
||||||
|
-- 20,430 root/engager blogs have any stored Posts rows at all (TL.db.md),
|
||||||
|
-- so most reblog/reply notes have no Posts row to check and must be kept,
|
||||||
|
-- not dropped by an inner join. COALESCE(p.IsActive, 1) = 1 keeps a note
|
||||||
|
-- unless its post is explicitly IsActive = 0 -- NULL (no Posts row, or a
|
||||||
|
-- stored row with no flag written) means live, per the schema's own
|
||||||
|
-- convention (TL.db.md, "Posts.IsActive and Notes.IsActive"). This is a
|
||||||
|
-- big filter in practice: of the blogs that qualified before it, most
|
||||||
|
-- have every one of their reblog/reply notes pointing at a since-removed
|
||||||
|
-- post, not just some -- verified against the live data, not assumed.
|
||||||
|
--
|
||||||
|
-- On DateAdded: this column is not written consistently -- most rows hold
|
||||||
|
-- ISO 'yyyy-MM-dd HH:mm:ss', but 17k+ hold US 'M/d/yy' from a 2025 bulk
|
||||||
|
-- import (see TL.db.md, "DateAdded is not written consistently"). As text,
|
||||||
|
-- those two shapes do not sort or compare against each other correctly, so
|
||||||
|
-- QUERY 0 normalises both to an ISO date before filtering. In the live data
|
||||||
|
-- every US-format row predates August 2026 anyway (only '12/23/25' and
|
||||||
|
-- '12/24/25' occur), so this makes no difference to the current answer --
|
||||||
|
-- it's here so the query stays correct if that ever changes.
|
||||||
|
-- ============================================================================
|
||||||
|
|
||||||
|
|
||||||
|
-- ----------------------------------------------------------------------------
|
||||||
|
-- QUERY 0 / THE ANSWER -- one row per qualifying blog, sorted by
|
||||||
|
-- earliest_reblog_or_reply_utc then DateAdded descending (normalised --
|
||||||
|
-- see the note above). No date cutoff, but scoped to HasBeenOutput = 0
|
||||||
|
-- AND IsActive = 1, and now excluding notes on a removed post (see the
|
||||||
|
-- header note above). 1,804 rows in the live data as of this revision --
|
||||||
|
-- down from 4,396 just before this exclusion was added, because most of
|
||||||
|
-- the blogs that dropped out had *every* reblog/reply note pointing at a
|
||||||
|
-- now-inactive post, not just some (the number moves between runs
|
||||||
|
-- regardless -- crawling and output flip HasBeenOutput/IsActive on live
|
||||||
|
-- rows).
|
||||||
|
--
|
||||||
|
-- earliest_reblog_or_reply_utc is the earliest TimeStamp among this
|
||||||
|
-- blog's reblog-or-reply notes (either type counts -- see the column
|
||||||
|
-- name); earliest_reblog_or_reply_postid and _root_blogid identify that
|
||||||
|
-- specific note's post: PostID + RootBlogId together, not PostID alone --
|
||||||
|
-- see TL.db.md ("345 post IDs exist under more than one blog"), same
|
||||||
|
-- caution as in find-notes-on-inactive-posts.sql. Resolve RootBlogId to a
|
||||||
|
-- name via Blogs (or Blogs, tolerating a miss) if you need it.
|
||||||
|
--
|
||||||
|
-- Getting "which note" rather than just "when" doesn't fit a plain
|
||||||
|
-- MIN()/GROUP BY -- an aggregate can tell you the earliest value but not
|
||||||
|
-- which row it came from. EarliestNote instead ranks each blog's
|
||||||
|
-- reblog/reply notes with ROW_NUMBER() OVER (PARTITION BY NoteBlogId
|
||||||
|
-- ORDER BY TimeStamp), and QUERY 0 takes rn = 1. The ORDER BY carries a
|
||||||
|
-- PostID tiebreak because (NoteBlogId, TimeStamp) is not unique in this
|
||||||
|
-- data -- ties exist (e.g. NoteBlogId 12 has 7 notes at the same
|
||||||
|
-- TimeStamp) -- so without a tiebreak the "earliest" pick would be
|
||||||
|
-- arbitrary among ties rather than deterministic.
|
||||||
|
--
|
||||||
|
-- EarliestNote also excludes notes on an inactive post before ranking
|
||||||
|
-- (see the header note above), so "earliest" means earliest surviving
|
||||||
|
-- note, not earliest overall -- a blog whose true-earliest note pointed
|
||||||
|
-- at a since-removed post now surfaces its next-earliest live one
|
||||||
|
-- instead. Applying the exclusion here, not as a filter on QUERY 0's
|
||||||
|
-- final rows, matters: filtering after ROW_NUMBER would have picked the
|
||||||
|
-- removed-post note as rn = 1 and then dropped the whole row instead of
|
||||||
|
-- promoting the next candidate.
|
||||||
|
-- ----------------------------------------------------------------------------
|
||||||
|
WITH BlogsSplit AS (
|
||||||
|
SELECT
|
||||||
|
b.BlogId,
|
||||||
|
b.BlogName,
|
||||||
|
b.DateAdded,
|
||||||
|
CASE WHEN b.DateAdded LIKE '____-__-__%' THEN 1 ELSE 0 END AS IsIso,
|
||||||
|
-- for the US 'M/d/yy' shape only: everything after the first '/'
|
||||||
|
substr(b.DateAdded, instr(b.DateAdded, '/') + 1) AS RestAfterMonth
|
||||||
|
FROM Blogs b
|
||||||
|
WHERE b.BlogId IS NOT NULL and HasBeenOutput = 0 and IsActive = 1 -- a blog can only match Notes if it has one
|
||||||
|
),
|
||||||
|
BlogsNorm AS (
|
||||||
|
SELECT
|
||||||
|
BlogId,
|
||||||
|
BlogName,
|
||||||
|
DateAdded,
|
||||||
|
CASE
|
||||||
|
WHEN IsIso = 1 THEN date(DateAdded)
|
||||||
|
ELSE date(
|
||||||
|
'20' || substr(RestAfterMonth, instr(RestAfterMonth, '/') + 1) || '-' ||
|
||||||
|
substr('00' || substr(DateAdded, 1, instr(DateAdded, '/') - 1), -2) || '-' ||
|
||||||
|
substr('00' || substr(RestAfterMonth, 1, instr(RestAfterMonth, '/') - 1), -2)
|
||||||
|
)
|
||||||
|
END AS DateAddedNorm
|
||||||
|
FROM BlogsSplit
|
||||||
|
),
|
||||||
|
EarliestNote AS (
|
||||||
|
SELECT
|
||||||
|
n.NoteBlogId,
|
||||||
|
n.RootBlogId,
|
||||||
|
n.PostID,
|
||||||
|
n.TimeStamp,
|
||||||
|
ROW_NUMBER() OVER (
|
||||||
|
PARTITION BY n.NoteBlogId
|
||||||
|
ORDER BY n.TimeStamp ASC, n.PostID ASC
|
||||||
|
) AS rn
|
||||||
|
FROM Notes n
|
||||||
|
JOIN NoteTypes t ON t.TypeId = n.TypeId
|
||||||
|
JOIN Blogs root_bn ON root_bn.BlogId = n.RootBlogId
|
||||||
|
LEFT JOIN Posts p ON p.BlogName = root_bn.BlogName
|
||||||
|
AND p.PostID = n.PostID
|
||||||
|
WHERE t.Type IN ('reblog')--, 'reply')
|
||||||
|
AND n.NoteBlogId IN (SELECT BlogId FROM BlogsNorm) -- scope the window to blogs we care about
|
||||||
|
AND COALESCE(p.IsActive, 1) = 1 -- exclude notes on a post explicitly marked removed
|
||||||
|
)
|
||||||
|
SELECT
|
||||||
|
bn.BlogId,
|
||||||
|
bn.BlogName || '.tumblr.com',
|
||||||
|
'''' || bn.blogname || ''',',
|
||||||
|
bn.DateAdded,
|
||||||
|
bn.DateAddedNorm,
|
||||||
|
datetime(en.TimeStamp, 'unixepoch') AS earliest_reblog_or_reply_utc,
|
||||||
|
en.PostID AS earliest_reblog_or_reply_postid,
|
||||||
|
en.RootBlogId AS earliest_reblog_or_reply_root_blogid
|
||||||
|
FROM BlogsNorm bn
|
||||||
|
JOIN EarliestNote en ON en.NoteBlogId = bn.BlogId AND en.rn = 1
|
||||||
|
ORDER BY earliest_reblog_or_reply_utc, bn.DateAddedNorm desc
|
||||||
|
limit 50;
|
||||||
|
|
||||||
|
|
||||||
|
-- ----------------------------------------------------------------------------
|
||||||
|
-- QUERY 1 -- same answer, with a per-blog breakdown of which type(s) fired
|
||||||
|
-- and how many. Useful once QUERY 0 has rows; redundant while it's empty.
|
||||||
|
-- ----------------------------------------------------------------------------
|
||||||
|
-- WITH BlogsSplit AS ( ... ), BlogsNorm AS ( ... ) -- reuse the CTEs above
|
||||||
|
--
|
||||||
|
-- SELECT
|
||||||
|
-- bn.BlogId,
|
||||||
|
-- bn.BlogName,
|
||||||
|
-- bn.DateAddedNorm,
|
||||||
|
-- SUM(CASE WHEN t.Type = 'reblog' THEN 1 ELSE 0 END) AS reblog_count,
|
||||||
|
-- SUM(CASE WHEN t.Type = 'reply' THEN 1 ELSE 0 END) AS reply_count
|
||||||
|
-- FROM BlogsNorm bn
|
||||||
|
-- JOIN Notes n ON n.NoteBlogId = bn.BlogId
|
||||||
|
-- JOIN NoteTypes t ON t.TypeId = n.TypeId
|
||||||
|
-- WHERE t.Type IN ('reblog', 'reply')
|
||||||
|
-- GROUP BY bn.BlogId, bn.BlogName, bn.DateAddedNorm
|
||||||
|
-- ORDER BY bn.DateAddedNorm;
|
||||||
|
|
||||||
|
|
||||||
|
-- ----------------------------------------------------------------------------
|
||||||
|
-- QUERY 2 -- put a date floor back, if wanted later.
|
||||||
|
-- Same as QUERY 0, with one extra line in the outer WHERE:
|
||||||
|
-- AND bn.DateAddedNorm >= '2026-09-01' -- or whatever cutoff
|
||||||
|
-- ----------------------------------------------------------------------------
|
||||||
|
</sql><current_tab id="0"/></tab_sql></sqlb_project>
|
||||||
|
|||||||
@@ -9,16 +9,22 @@ WHERE
|
|||||||
ORDER BY
|
ORDER BY
|
||||||
postdate desc</sql><sql name="SQL 2*">SELECT
|
postdate desc</sql><sql name="SQL 2*">SELECT
|
||||||
datetime(TimeStamp, 'unixepoch'),
|
datetime(TimeStamp, 'unixepoch'),
|
||||||
RootBlogName || '.tumblr.com/post/' || N.postid,
|
rbn.BlogName || '.tumblr.com/post/' || N.postid,
|
||||||
*,
|
*,
|
||||||
NoteBlogName || '.tumblr.com'
|
nbn.BlogName || '.tumblr.com'
|
||||||
FROM
|
FROM
|
||||||
Notes N
|
Notes N
|
||||||
inner JOIN
|
inner JOIN
|
||||||
Posts P on P.PostID = N.PostID and P.BlogName = N.RootBlogName
|
Blogs rbn on rbn.BlogId = N.RootBlogId
|
||||||
WHERE RootBlogName NOT IN ('xlittle-ghost', 'glimmerin-darlin', 'vvenus-child')␍
|
inner JOIN
|
||||||
and type like 'r%'␍
|
Blogs nbn on nbn.BlogId = N.NoteBlogId
|
||||||
and RootBlogName = 'zomb-eh'␍
|
inner JOIN
|
||||||
|
NoteTypes t on t.TypeId = N.TypeId
|
||||||
|
inner JOIN
|
||||||
|
Posts P on P.PostID = N.PostID and P.BlogName = rbn.BlogName
|
||||||
|
WHERE rbn.BlogName NOT IN ('xlittle-ghost', 'glimmerin-darlin', 'vvenus-child')
|
||||||
|
and t.Type like 'r%'
|
||||||
|
and rbn.BlogName = 'zomb-eh'
|
||||||
and P.HasImage = 1
|
and P.HasImage = 1
|
||||||
ORDER BY
|
ORDER BY
|
||||||
TimeStamp desc</sql><current_tab id="1"/></tab_sql></sqlb_project>
|
TimeStamp desc</sql><current_tab id="1"/></tab_sql></sqlb_project>
|
||||||
|
|||||||
@@ -216,18 +216,19 @@ namespace URLNotesGrabberCORE
|
|||||||
#region Notes integer schema
|
#region Notes integer schema
|
||||||
|
|
||||||
// Notes stopped storing names on 2026-08-07: RootBlogName/NoteBlogName/Type became
|
// Notes stopped storing names on 2026-08-07: RootBlogName/NoteBlogName/Type became
|
||||||
// RootBlogId/NoteBlogId/TypeId, resolved through BlogNames and NoteTypes. There is no
|
// RootBlogId/NoteBlogId/TypeId, resolved through Blogs.BlogId and NoteTypes. There is no
|
||||||
// compatibility view -- a query naming an old column fails outright, so this is a hard
|
// compatibility view -- a query naming an old column fails outright, so this is a hard
|
||||||
// cut rather than an optional column like IsActive. See TL.db.md.
|
// cut rather than an optional column like IsActive. See TL.db.md.
|
||||||
//
|
//
|
||||||
|
// Blogs.BlogId is the only ID authority as of 2026-09-28; the BlogNames table that used
|
||||||
|
// to hold the IDs is gone. Every name in Notes has a Blogs row, created by RegisterBlog.
|
||||||
|
//
|
||||||
// Two shapes recur below and are spelled out inline rather than hidden behind a helper,
|
// Two shapes recur below and are spelled out inline rather than hidden behind a helper,
|
||||||
// so that every statement reads as the SQL it actually runs:
|
// so that every statement reads as the SQL it actually runs:
|
||||||
// (SELECT BlogId FROM BlogNames WHERE BlogName = @name) -- unique-index probe, 20k rows
|
// (SELECT BlogId FROM Blogs WHERE BlogName = @name) -- primary-key probe
|
||||||
// (SELECT TypeId FROM NoteTypes WHERE Type = 'reply') -- 5 rows, effectively free
|
// (SELECT TypeId FROM NoteTypes WHERE Type = 'reply') -- 5 rows, effectively free
|
||||||
// Joining Notes to Blogs is the one case that must NOT route through BlogNames: Blogs
|
// Joining Notes to Blogs is N.NoteBlogId = B.BlogId, a single hop on the unique index.
|
||||||
// carries its own BlogId, so N.NoteBlogId = B.BlogId is a single integer hop. Joining
|
// Joining Notes to Posts goes through Blogs too -- Posts has only BlogName.
|
||||||
// Notes to Posts is the opposite case -- Posts has only BlogName, so it has to go
|
|
||||||
// through BlogNames.
|
|
||||||
|
|
||||||
/// <summary>
|
/// <summary>
|
||||||
/// True when the exception is a duplicate-key collision on Notes. The message embeds the
|
/// True when the exception is a duplicate-key collision on Notes. The message embeds the
|
||||||
@@ -242,14 +243,30 @@ namespace URLNotesGrabberCORE
|
|||||||
}
|
}
|
||||||
|
|
||||||
/// <summary>
|
/// <summary>
|
||||||
/// Gives a blog name an ID if it does not have one. No read-back and no round trip -- a
|
/// Ensures a blog has a Blogs row and a BlogId, so a note can point at it. No read-back
|
||||||
/// name that is already registered keeps the ID that 1.18M Notes rows point at.
|
/// and no round trip -- a blog that already has an ID keeps the one Notes rows point at.
|
||||||
|
/// Unlike AddBlog this does not skip "deact" names: a note by a deactivated blog still
|
||||||
|
/// needs an ID, and Blogs is the only place one can live.
|
||||||
|
/// Assigning the ID is bookkeeping, not a content change, so DateModified is not touched.
|
||||||
|
/// MAX(BlogId) + 1 cannot hand out a used ID because trg_Blogs_BlogId_NoDelete stops any
|
||||||
|
/// row that holds one from being deleted.
|
||||||
/// </summary>
|
/// </summary>
|
||||||
private static void RegisterBlogName(SQLiteConnection connection, SQLiteTransaction? transaction, string blogName)
|
private static void RegisterBlog(SQLiteConnection connection, SQLiteTransaction? transaction, string blogName)
|
||||||
{
|
{
|
||||||
using SQLiteCommand command = new SQLiteCommand("INSERT OR IGNORE INTO BlogNames (BlogName) VALUES (@BlogName)", connection, transaction);
|
string now = DateTime.Now.ToString("yyyy-MM-dd HH:mm:ss");
|
||||||
command.Parameters.AddWithValue("@BlogName", blogName);
|
|
||||||
command.ExecuteNonQuery();
|
using (SQLiteCommand command = new SQLiteCommand("INSERT OR IGNORE INTO Blogs (BlogName, DateAdded, DateModified, DateCreated) VALUES (@BlogName, @Now, @Now, @Now)", connection, transaction))
|
||||||
|
{
|
||||||
|
command.Parameters.AddWithValue("@BlogName", blogName);
|
||||||
|
command.Parameters.AddWithValue("@Now", now);
|
||||||
|
command.ExecuteNonQuery();
|
||||||
|
}
|
||||||
|
|
||||||
|
using (SQLiteCommand command = new SQLiteCommand("UPDATE Blogs SET BlogId = (SELECT IFNULL(MAX(BlogId), 0) + 1 FROM Blogs) WHERE BlogName = @BlogName AND BlogId IS NULL", connection, transaction))
|
||||||
|
{
|
||||||
|
command.Parameters.AddWithValue("@BlogName", blogName);
|
||||||
|
command.ExecuteNonQuery();
|
||||||
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
/// <summary>
|
/// <summary>
|
||||||
@@ -590,12 +607,17 @@ namespace URLNotesGrabberCORE
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
public static void AddPost(string blogName, long postID, string reblogURL, string postDate, string postURL, string slug, string reblogKey, string reblogName, string summary, string quote, string body, string tags, string link, string photoURL, string photoCaption, string downloadedFiles, string audioCaption, string question, string answer, string title, bool hasImage, bool byLikes = false, string? DBPath = null, string? rootBlogName = null, string? rootURL = null)
|
// postType: a canonical PostTypes name, or null when the caller has no trustworthy type.
|
||||||
|
// Null is stored as NULL rather than guessed at -- OutputMode skips untyped rows, so a
|
||||||
|
// null costs one export line, whereas a wrong value would create a wrongly named file.
|
||||||
|
public static void AddPost(string blogName, long postID, string reblogURL, string postDate, string postURL, string slug, string reblogKey, string reblogName, string summary, string quote, string body, string tags, string link, string photoURL, string photoCaption, string downloadedFiles, string audioCaption, string question, string answer, string title, bool hasImage, bool byLikes = false, string? DBPath = null, string? rootBlogName = null, string? rootURL = null, string? postType = null)
|
||||||
{
|
{
|
||||||
DBPath ??= GetDefaultDbPath();
|
DBPath ??= GetDefaultDbPath();
|
||||||
|
postType = PostTypes.Normalize(postType);
|
||||||
|
postDate = PostDates.Normalize(postDate)!;
|
||||||
try { AddBlog(blogName, byLikes, DBPath); } catch { }
|
try { AddBlog(blogName, byLikes, DBPath); } catch { }
|
||||||
try { UpdatePostSetDate(blogName, postID, postDate, DBPath); } catch { }
|
try { UpdatePostSetDate(blogName, postID, postDate, DBPath); } catch { }
|
||||||
try { UpdatePost(blogName, postID, reblogURL, postDate, postURL, slug, reblogKey, reblogName, summary, quote, body, tags, link, photoURL, photoCaption, downloadedFiles, audioCaption, question, answer, title, hasImage, byLikes, DBPath, rootBlogName, rootURL); } catch { }
|
try { UpdatePost(blogName, postID, reblogURL, postDate, postURL, slug, reblogKey, reblogName, summary, quote, body, tags, link, photoURL, photoCaption, downloadedFiles, audioCaption, question, answer, title, hasImage, byLikes, DBPath, rootBlogName, rootURL, postType); } catch { }
|
||||||
|
|
||||||
SQLiteConnection connection;
|
SQLiteConnection connection;
|
||||||
bool ownsConnection;
|
bool ownsConnection;
|
||||||
@@ -637,9 +659,10 @@ namespace URLNotesGrabberCORE
|
|||||||
RootBlogName,
|
RootBlogName,
|
||||||
RootURL,
|
RootURL,
|
||||||
HasImage,
|
HasImage,
|
||||||
ByLikes
|
ByLikes,
|
||||||
|
PostType
|
||||||
) VALUES (" +
|
) VALUES (" +
|
||||||
Q(blogName) + ", " + postID + ", " + Q(reblogURL) + ", " + Q(postDate) + ", " + Q(postURL) + ", " + Q(slug) + ", " + Q(reblogKey) + ", " + Q(reblogName) + ", " + Q(summary) + ", " + Q(quote) + ", " + Q(body) + ", " + Q(tags) + ", " + Q(link) + ", " + Q(photoURL) + ", " + Q(photoCaption) + ", " + Q(downloadedFiles) + ", " + Q(audioCaption) + ", " + Q(question) + ", " + Q(answer) + ", " + Q(title) + ", " + Q(DateTime.Now.ToString("yyyy-MM-dd HH:mm:ss")) + ", " + Q(DateTime.Now.ToString("yyyy-MM-dd HH:mm:ss")) + ", " + Q(rootBlogName ?? ".") + ", " + Q(rootURL ?? ".") + ", " + (hasImage ? 1 : 0) + ", " + (byLikes ? 1 : 0) + ")";
|
Q(blogName) + ", " + postID + ", " + Q(reblogURL) + ", " + Q(postDate) + ", " + Q(postURL) + ", " + Q(slug) + ", " + Q(reblogKey) + ", " + Q(reblogName) + ", " + Q(summary) + ", " + Q(quote) + ", " + Q(body) + ", " + Q(tags) + ", " + Q(link) + ", " + Q(photoURL) + ", " + Q(photoCaption) + ", " + Q(downloadedFiles) + ", " + Q(audioCaption) + ", " + Q(question) + ", " + Q(answer) + ", " + Q(title) + ", " + Q(DateTime.Now.ToString("yyyy-MM-dd HH:mm:ss")) + ", " + Q(DateTime.Now.ToString("yyyy-MM-dd HH:mm:ss")) + ", " + Q(rootBlogName ?? ".") + ", " + Q(rootURL ?? ".") + ", " + (hasImage ? 1 : 0) + ", " + (byLikes ? 1 : 0) + ", " + (postType == null ? "NULL" : Q(postType)) + ")";
|
||||||
SQLiteCommand command = new SQLiteCommand(sql, connection);
|
SQLiteCommand command = new SQLiteCommand(sql, connection);
|
||||||
|
|
||||||
int rowsInserted = 0;
|
int rowsInserted = 0;
|
||||||
@@ -756,7 +779,6 @@ namespace URLNotesGrabberCORE
|
|||||||
{
|
{
|
||||||
DBPath ??= GetDefaultDbPath();
|
DBPath ??= GetDefaultDbPath();
|
||||||
//try { AddPost(rootBlogName, postID, DBPath); } catch { }
|
//try { AddPost(rootBlogName, postID, DBPath); } catch { }
|
||||||
try { AddBlog(noteBlogName, false, DBPath); } catch { }
|
|
||||||
|
|
||||||
using SQLiteConnection connection2 = new SQLiteConnection("Data Source=" + DBPath);
|
using SQLiteConnection connection2 = new SQLiteConnection("Data Source=" + DBPath);
|
||||||
int rowsInserted = 0;
|
int rowsInserted = 0;
|
||||||
@@ -765,23 +787,24 @@ namespace URLNotesGrabberCORE
|
|||||||
{
|
{
|
||||||
connection2.Open();
|
connection2.Open();
|
||||||
|
|
||||||
// Notes stores integer IDs, so both participants and the type have to exist in
|
// Notes stores integer IDs, so both participants need a Blogs row with a BlogId,
|
||||||
// their lookup table before the note can point at them.
|
// and the type a NoteTypes row, before the note can point at them. RegisterBlog
|
||||||
|
// also does what the AddBlog call here used to: register the note's blog.
|
||||||
//
|
//
|
||||||
// All four statements run in one transaction so a crash cannot leave a name or a
|
// All statements run in one transaction so a crash cannot leave a blog or a
|
||||||
// type registered with no note. The transaction is committed before the console
|
// type registered with no note. The transaction is committed before the console
|
||||||
// output below, which sleeps -- a write lock must not be held across that.
|
// output below, which sleeps -- a write lock must not be held across that.
|
||||||
using (SQLiteTransaction transaction = connection2.BeginTransaction())
|
using (SQLiteTransaction transaction = connection2.BeginTransaction())
|
||||||
{
|
{
|
||||||
RegisterBlogName(connection2, transaction, rootBlogName);
|
RegisterBlog(connection2, transaction, rootBlogName);
|
||||||
RegisterBlogName(connection2, transaction, noteBlogName);
|
RegisterBlog(connection2, transaction, noteBlogName);
|
||||||
RegisterNoteType(connection2, transaction, type ?? string.Empty);
|
RegisterNoteType(connection2, transaction, type ?? string.Empty);
|
||||||
|
|
||||||
// INSERT OR IGNORE, and no IsActive in the column list: re-crawling a note
|
// INSERT OR IGNORE, and no IsActive in the column list: re-crawling a note
|
||||||
// that was removed elsewhere leaves the existing row -- and its flag -- alone.
|
// that was removed elsewhere leaves the existing row -- and its flag -- alone.
|
||||||
string sql = "INSERT OR IGNORE INTO Notes (RootBlogId, NoteBlogId, PostID, TimeStamp, TypeId, DatetimeCrawled, DateModified, DateCreated) " +
|
string sql = "INSERT OR IGNORE INTO Notes (RootBlogId, NoteBlogId, PostID, TimeStamp, TypeId, DatetimeCrawled, DateModified, DateCreated) " +
|
||||||
"SELECT (SELECT BlogId FROM BlogNames WHERE BlogName = @rootBlogName), " +
|
"SELECT (SELECT BlogId FROM Blogs WHERE BlogName = @rootBlogName), " +
|
||||||
" (SELECT BlogId FROM BlogNames WHERE BlogName = @noteBlogName), " +
|
" (SELECT BlogId FROM Blogs WHERE BlogName = @noteBlogName), " +
|
||||||
" @PostID, @TimeStamp, " +
|
" @PostID, @TimeStamp, " +
|
||||||
" (SELECT TypeId FROM NoteTypes WHERE Type = @Type), " +
|
" (SELECT TypeId FROM NoteTypes WHERE Type = @Type), " +
|
||||||
" @DatetimeCrawled, @DateModified, @DateCreated";
|
" @DatetimeCrawled, @DateModified, @DateCreated";
|
||||||
@@ -809,7 +832,7 @@ namespace URLNotesGrabberCORE
|
|||||||
Console.ForegroundColor = ConsoleColor.Green;
|
Console.ForegroundColor = ConsoleColor.Green;
|
||||||
Console.WriteLine("{2}\t{0}\t{1}", UnixTimeStampToDateTime(timestamp), noteBlogName, type);
|
Console.WriteLine("{2}\t{0}\t{1}", UnixTimeStampToDateTime(timestamp), noteBlogName, type);
|
||||||
Console.ForegroundColor = previousColor;
|
Console.ForegroundColor = previousColor;
|
||||||
Thread.Sleep(250); // Brief pause to make new notes more noticeable in the console output
|
Thread.Sleep(125); // Brief pause to make new notes more noticeable in the console output
|
||||||
}
|
}
|
||||||
else
|
else
|
||||||
{
|
{
|
||||||
@@ -853,9 +876,12 @@ namespace URLNotesGrabberCORE
|
|||||||
///
|
///
|
||||||
/// </summary>
|
/// </summary>
|
||||||
/// <param name="withoutNotesOnly"></param>
|
/// <param name="withoutNotesOnly"></param>
|
||||||
|
/// <param name="ignoreRefreshCooldown">Drops the age gate on the periodic re-queue branch (--force).</param>
|
||||||
|
/// <param name="fromDate">Lower bound on the *original post's* PostDate for the periodic re-queue branch (--fromDate). Independent of ignoreRefreshCooldown -- applies whether or not --force is also given.</param>
|
||||||
|
/// <param name="toDate">Upper bound on the *original post's* PostDate for the periodic re-queue branch (--toDate). Same independence from ignoreRefreshCooldown as fromDate.</param>
|
||||||
/// <param name="DBPath"></param>
|
/// <param name="DBPath"></param>
|
||||||
/// <returns>blogName, postID, lastNoteTimestamp, notesGatheredTimestamp</returns>
|
/// <returns>blogName, postID, lastNoteTimestamp, notesGatheredTimestamp</returns>
|
||||||
public static List<Tuple<string, long, long, long>> GetPosts(bool withoutNotesOnly = false, DateTime? beforeDate = null, string? blogName = null, string? DBPath = null)
|
public static List<Tuple<string, long, long, long>> GetPosts(bool withoutNotesOnly = false, DateTime? beforeDate = null, string? blogName = null, bool ignoreRefreshCooldown = false, DateTime? fromDate = null, DateTime? toDate = null, string? DBPath = null)
|
||||||
{
|
{
|
||||||
DBPath ??= GetDefaultDbPath();
|
DBPath ??= GetDefaultDbPath();
|
||||||
using SQLiteConnection connection = new SQLiteConnection("Data Source=" + DBPath);
|
using SQLiteConnection connection = new SQLiteConnection("Data Source=" + DBPath);
|
||||||
@@ -880,16 +906,48 @@ namespace URLNotesGrabberCORE
|
|||||||
beforeDateFilter = $"WHERE (U.NotesGatheredDateTime < {unixTimestamp} OR U.NotesGatheredDateTime IS NULL)" + Environment.NewLine;
|
beforeDateFilter = $"WHERE (U.NotesGatheredDateTime < {unixTimestamp} OR U.NotesGatheredDateTime IS NULL)" + Environment.NewLine;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// Blogs.IsActive is the crawler's work-selection flag (Rolodex removal sets it to 0)
|
||||||
|
// and is independent of Posts.IsActive/Notes.IsActive -- deactivating a blog does not
|
||||||
|
// touch its posts' own IsActive column. AndIsActive("Posts", ...) above therefore does
|
||||||
|
// not catch a deactivated blog; this join against the source rows is what does, so a
|
||||||
|
// blog taken IsActive = 0 in Blogs stops being re-queued by --collect 1 even if its
|
||||||
|
// posts were never individually marked inactive. Blogs.BlogName is that table's PRIMARY
|
||||||
|
// KEY, so the join rides an index rather than scanning it.
|
||||||
|
//
|
||||||
// Same WHERE/AND juggling WhereIsActive does, extended to the optional blog predicate:
|
// Same WHERE/AND juggling WhereIsActive does, extended to the optional blog predicate:
|
||||||
// either clause may be absent, so the first one present has to open the WHERE.
|
// either clause may be absent, so the first one present has to open the WHERE.
|
||||||
string sourceClause = AndIsActive("Posts", "P", DBPath) + (filterByBlog ? " AND P.BlogName = @blogName" : string.Empty);
|
string sourceClause = AndIsActive("Posts", "P", DBPath) + " AND COALESCE(BL.IsActive, 1) = 1" + (filterByBlog ? " AND P.BlogName = @blogName" : string.Empty);
|
||||||
string sourceFilter = sourceClause.Length == 0 ? string.Empty : " WHERE" + sourceClause.Substring(" AND".Length);
|
string sourceFilter = sourceClause.Length == 0 ? string.Empty : " WHERE" + sourceClause.Substring(" AND".Length);
|
||||||
|
|
||||||
// The zomb-eh branch re-queues that blog's *already collected* posts every 3 days. It needs
|
// The zomb-eh branch re-queues that blog's *already collected* posts every 3 days. It needs
|
||||||
// no blog-filter handling of its own: it reads PostsWithCount, which the filter has already
|
// no blog-filter or IsActive handling of its own: it reads PostsWithCount, which the source
|
||||||
// scoped, so it contributes its rows when the filter names zomb-eh and nothing otherwise.
|
// filter above -- Blogs.IsActive included -- has already scoped, so it contributes its rows
|
||||||
// That keeps a filtered worklist a strict subset of the unfiltered one -- "--collect 1 X"
|
// only when zomb-eh itself is still IsActive = 1 there. That keeps a filtered worklist a
|
||||||
// returns exactly the rows "--collect 1" would have returned for X.
|
// strict subset of the unfiltered one -- "--collect 1 X" returns exactly the rows
|
||||||
|
// "--collect 1" would have returned for X.
|
||||||
|
//
|
||||||
|
// --force drops the age gate only. NotFound = 0 and the IsActive/blog scoping above still
|
||||||
|
// apply: the flag is "re-collect early", not "collect rows every other path excludes".
|
||||||
|
string refreshCooldownClause = ignoreRefreshCooldown
|
||||||
|
? string.Empty
|
||||||
|
: " AND NotesGatheredDateTime < unixepoch('now', 'localtime', '-3 days')" + Environment.NewLine;
|
||||||
|
|
||||||
|
// --fromDate bounds the *original post's* PostDate, not the re-collect cooldown --
|
||||||
|
// it stacks with refreshCooldownClause instead of replacing it, so it applies the
|
||||||
|
// same way whether or not --force also dropped the cooldown. A NULL PostDate never
|
||||||
|
// satisfies ">=" and is excluded, same as an unfiltered run would still include it
|
||||||
|
// (there's nothing to compare here, so this only narrows, never widens, the result).
|
||||||
|
string fromDateClause = fromDate.HasValue
|
||||||
|
? " AND PostDate >= @fromDate" + Environment.NewLine
|
||||||
|
: string.Empty;
|
||||||
|
|
||||||
|
// --toDate is the same deal, mirrored: stacks alongside fromDateClause/
|
||||||
|
// refreshCooldownClause rather than replacing either, so --fromDate and --toDate
|
||||||
|
// can be given together (or alone) and both hold with or without --force.
|
||||||
|
string toDateClause = toDate.HasValue
|
||||||
|
? " AND PostDate <= @toDate" + Environment.NewLine
|
||||||
|
: string.Empty;
|
||||||
|
|
||||||
string refreshBranch =
|
string refreshBranch =
|
||||||
"" + Environment.NewLine +
|
"" + Environment.NewLine +
|
||||||
" UNION " + Environment.NewLine +
|
" UNION " + Environment.NewLine +
|
||||||
@@ -904,7 +962,9 @@ namespace URLNotesGrabberCORE
|
|||||||
" FROM PostsWithCount" + Environment.NewLine +
|
" FROM PostsWithCount" + Environment.NewLine +
|
||||||
" WHERE BlogName = 'zomb-eh'" + Environment.NewLine +
|
" WHERE BlogName = 'zomb-eh'" + Environment.NewLine +
|
||||||
" AND NotFound = 0" + Environment.NewLine +
|
" AND NotFound = 0" + Environment.NewLine +
|
||||||
" AND NotesGatheredDateTime < unixepoch('now', 'localtime', '-3 days')" + Environment.NewLine;
|
refreshCooldownClause +
|
||||||
|
fromDateClause +
|
||||||
|
toDateClause;
|
||||||
|
|
||||||
sql = "WITH PostsWithCount AS" + Environment.NewLine +
|
sql = "WITH PostsWithCount AS" + Environment.NewLine +
|
||||||
"(" + Environment.NewLine +
|
"(" + Environment.NewLine +
|
||||||
@@ -917,7 +977,7 @@ namespace URLNotesGrabberCORE
|
|||||||
" P.HasNotesGathered," + Environment.NewLine +
|
" P.HasNotesGathered," + Environment.NewLine +
|
||||||
" P.NotFound," + Environment.NewLine +
|
" P.NotFound," + Environment.NewLine +
|
||||||
" P.PostDate" + Environment.NewLine +
|
" P.PostDate" + Environment.NewLine +
|
||||||
" FROM Posts P" + sourceFilter + Environment.NewLine +
|
" FROM Posts P LEFT JOIN Blogs BL ON BL.BlogName = P.BlogName" + sourceFilter + Environment.NewLine +
|
||||||
")," + Environment.NewLine +
|
")," + Environment.NewLine +
|
||||||
"Unioned AS" + Environment.NewLine +
|
"Unioned AS" + Environment.NewLine +
|
||||||
"(" + Environment.NewLine +
|
"(" + Environment.NewLine +
|
||||||
@@ -988,6 +1048,13 @@ namespace URLNotesGrabberCORE
|
|||||||
if (filterByBlog)
|
if (filterByBlog)
|
||||||
command.Parameters.AddWithValue("@blogName", blogName);
|
command.Parameters.AddWithValue("@blogName", blogName);
|
||||||
|
|
||||||
|
// Only ever referenced by the zomb-eh refresh branch, which only exists when
|
||||||
|
// withoutNotesOnly is true -- harmless to bind unconditionally otherwise.
|
||||||
|
if (fromDate.HasValue)
|
||||||
|
command.Parameters.AddWithValue("@fromDate", fromDate.Value.ToString("yyyy-MM-dd HH:mm:ss"));
|
||||||
|
if (toDate.HasValue)
|
||||||
|
command.Parameters.AddWithValue("@toDate", toDate.Value.ToString("yyyy-MM-dd HH:mm:ss"));
|
||||||
|
|
||||||
using (SQLiteDataReader reader = command.ExecuteReader())
|
using (SQLiteDataReader reader = command.ExecuteReader())
|
||||||
{
|
{
|
||||||
while (reader.Read())
|
while (reader.Read())
|
||||||
@@ -1076,11 +1143,11 @@ namespace URLNotesGrabberCORE
|
|||||||
{
|
{
|
||||||
connection.Open();
|
connection.Open();
|
||||||
|
|
||||||
string sql = "SELECT DISTINCT BN.BlogName as blogName, N.PostID" +
|
string sql = "SELECT DISTINCT RB.BlogName as blogName, N.PostID" +
|
||||||
" FROM Notes N" +
|
" FROM Notes N" +
|
||||||
" INNER JOIN BlogNames BN ON BN.BlogId = N.RootBlogId" +
|
" INNER JOIN Blogs RB ON RB.BlogId = N.RootBlogId" +
|
||||||
" WHERE N.TypeId = (SELECT TypeId FROM NoteTypes WHERE Type = 'reply')" + AndIsActive("Notes", "N", DBPath) +
|
" WHERE N.TypeId = (SELECT TypeId FROM NoteTypes WHERE Type = 'reply')" + AndIsActive("Notes", "N", DBPath) +
|
||||||
" ORDER BY BN.BlogName, N.PostID";
|
" ORDER BY RB.BlogName, N.PostID";
|
||||||
|
|
||||||
using (SQLiteCommand command = new SQLiteCommand(sql, connection))
|
using (SQLiteCommand command = new SQLiteCommand(sql, connection))
|
||||||
{
|
{
|
||||||
@@ -1121,11 +1188,11 @@ namespace URLNotesGrabberCORE
|
|||||||
connection.Open();
|
connection.Open();
|
||||||
|
|
||||||
// Grouped on the integer rather than the name: the group key is what gets sorted,
|
// Grouped on the integer rather than the name: the group key is what gets sorted,
|
||||||
// and BN.BlogName comes along for free off the join.
|
// and RB.BlogName comes along for free off the join.
|
||||||
string sql = @"SELECT BN.BlogName as blogName, N.PostID,
|
string sql = @"SELECT RB.BlogName as blogName, N.PostID,
|
||||||
MAX(N.TimeStamp) as LatestTimestamp
|
MAX(N.TimeStamp) as LatestTimestamp
|
||||||
FROM Notes N
|
FROM Notes N
|
||||||
INNER JOIN BlogNames BN ON BN.BlogId = N.RootBlogId
|
INNER JOIN Blogs RB ON RB.BlogId = N.RootBlogId
|
||||||
WHERE N.TypeId = (SELECT TypeId FROM NoteTypes WHERE Type = 'reply')
|
WHERE N.TypeId = (SELECT TypeId FROM NoteTypes WHERE Type = 'reply')
|
||||||
AND (N.replyText IS NULL OR N.replyText = '' OR N.replyText = '.')" + AndIsActive("Notes", "N", DBPath) + @"
|
AND (N.replyText IS NULL OR N.replyText = '' OR N.replyText = '.')" + AndIsActive("Notes", "N", DBPath) + @"
|
||||||
GROUP BY N.RootBlogId, N.PostID
|
GROUP BY N.RootBlogId, N.PostID
|
||||||
@@ -1170,13 +1237,13 @@ namespace URLNotesGrabberCORE
|
|||||||
{
|
{
|
||||||
connection.Open();
|
connection.Open();
|
||||||
|
|
||||||
// Posts carries only BlogName, so this is the one join to Notes that has to go
|
// Posts carries only BlogName, so the join to Notes goes through Blogs -- there is
|
||||||
// through BlogNames -- there is no Posts.BlogId to hop on. The name predicate is
|
// no Posts.BlogId to hop on. Each post's name is a primary-key probe on Blogs,
|
||||||
// pushed into the 20k-row lookup, which then feeds integers to the Notes key.
|
// which then feeds an integer to the Notes key.
|
||||||
string sql = @"SELECT DISTINCT P.BlogName, P.PostID, MAX(N.TimeStamp) as LatestTimestamp
|
string sql = @"SELECT DISTINCT P.BlogName, P.PostID, MAX(N.TimeStamp) as LatestTimestamp
|
||||||
FROM Posts P
|
FROM Posts P
|
||||||
INNER JOIN BlogNames RBN ON RBN.BlogName = P.BlogName
|
INNER JOIN Blogs RB ON RB.BlogName = P.BlogName
|
||||||
INNER JOIN Notes N ON N.RootBlogId = RBN.BlogId AND N.PostID = P.PostID
|
INNER JOIN Notes N ON N.RootBlogId = RB.BlogId AND N.PostID = P.PostID
|
||||||
WHERE P.NotFound = 0
|
WHERE P.NotFound = 0
|
||||||
AND N.TypeId = (SELECT TypeId FROM NoteTypes WHERE Type = 'reply')
|
AND N.TypeId = (SELECT TypeId FROM NoteTypes WHERE Type = 'reply')
|
||||||
AND (N.replyText IS NULL OR N.replyText = '' OR N.replyText = '.')" + AndIsActive("Posts", "P", DBPath) + AndIsActive("Notes", "N", DBPath) + @"
|
AND (N.replyText IS NULL OR N.replyText = '' OR N.replyText = '.')" + AndIsActive("Posts", "P", DBPath) + AndIsActive("Notes", "N", DBPath) + @"
|
||||||
@@ -1365,8 +1432,7 @@ namespace URLNotesGrabberCORE
|
|||||||
try
|
try
|
||||||
{
|
{
|
||||||
connection.Open();
|
connection.Open();
|
||||||
// Blogs is reached in one integer hop off Blogs.BlogId, not through BlogNames --
|
// Blogs is reached in one integer hop off Blogs.BlogId.
|
||||||
// that would add a hop and end in the text comparison the migration removed.
|
|
||||||
// The negated form is only correct because Notes.TypeId is NOT NULL.
|
// The negated form is only correct because Notes.TypeId is NOT NULL.
|
||||||
string sql = "";
|
string sql = "";
|
||||||
if (reblogsOnly)
|
if (reblogsOnly)
|
||||||
@@ -1653,8 +1719,8 @@ namespace URLNotesGrabberCORE
|
|||||||
connection.Open();
|
connection.Open();
|
||||||
|
|
||||||
string sql = "UPDATE Notes SET TimeStamp = @timestamp, DateModified = @dateModified " +
|
string sql = "UPDATE Notes SET TimeStamp = @timestamp, DateModified = @dateModified " +
|
||||||
"WHERE RootBlogId = (SELECT BlogId FROM BlogNames WHERE BlogName = @rootBlogName) " +
|
"WHERE RootBlogId = (SELECT BlogId FROM Blogs WHERE BlogName = @rootBlogName) " +
|
||||||
"AND NoteBlogId = (SELECT BlogId FROM BlogNames WHERE BlogName = @noteBlogName) " +
|
"AND NoteBlogId = (SELECT BlogId FROM Blogs WHERE BlogName = @noteBlogName) " +
|
||||||
"AND PostID = @postID AND IFNULL(TimeStamp, 0) <> @timestamp";
|
"AND PostID = @postID AND IFNULL(TimeStamp, 0) <> @timestamp";
|
||||||
using (SQLiteCommand command = new SQLiteCommand(sql, connection))
|
using (SQLiteCommand command = new SQLiteCommand(sql, connection))
|
||||||
{
|
{
|
||||||
@@ -1679,9 +1745,10 @@ namespace URLNotesGrabberCORE
|
|||||||
return false;
|
return false;
|
||||||
}
|
}
|
||||||
|
|
||||||
public static void UpdatePost(string blogName, long postID, string reblogURL, string postDate, string postURL, string slug, string reblogKey, string reblogName, string summary, string quote, string body, string tags, string link, string photoURL, string photoCaption, string downloadedFiles, string audioCaption, string question, string answer, string title, bool hasImage, bool byLikes = false, string? DBPath = null, string? rootBlogName = null, string? rootURL = null)
|
public static void UpdatePost(string blogName, long postID, string reblogURL, string postDate, string postURL, string slug, string reblogKey, string reblogName, string summary, string quote, string body, string tags, string link, string photoURL, string photoCaption, string downloadedFiles, string audioCaption, string question, string answer, string title, bool hasImage, bool byLikes = false, string? DBPath = null, string? rootBlogName = null, string? rootURL = null, string? postType = null)
|
||||||
{
|
{
|
||||||
DBPath ??= GetDefaultDbPath();
|
DBPath ??= GetDefaultDbPath();
|
||||||
|
postType = PostTypes.Normalize(postType);
|
||||||
|
|
||||||
SQLiteConnection connection;
|
SQLiteConnection connection;
|
||||||
bool ownsConnection;
|
bool ownsConnection;
|
||||||
@@ -1729,6 +1796,10 @@ namespace URLNotesGrabberCORE
|
|||||||
sql += "RootBlogName = CASE WHEN @rootBlogName IS NULL OR @rootBlogName = '' OR @rootBlogName = '.' THEN RootBlogName ELSE @rootBlogName END, ";
|
sql += "RootBlogName = CASE WHEN @rootBlogName IS NULL OR @rootBlogName = '' OR @rootBlogName = '.' THEN RootBlogName ELSE @rootBlogName END, ";
|
||||||
sql += "RootURL = CASE WHEN @rootURL IS NULL OR @rootURL = '' OR @rootURL = '.' THEN RootURL ELSE @rootURL END, ";
|
sql += "RootURL = CASE WHEN @rootURL IS NULL OR @rootURL = '' OR @rootURL = '.' THEN RootURL ELSE @rootURL END, ";
|
||||||
sql += "hasImage = @hasImage, ";
|
sql += "hasImage = @hasImage, ";
|
||||||
|
// Fill in a missing type, never overwrite one. A type derived by --ingest from a
|
||||||
|
// real export filename is authoritative; this path's type is only as good as the
|
||||||
|
// folder it was crawled from, so it must not win over an existing value.
|
||||||
|
sql += "PostType = IFNULL(PostType, @postType), ";
|
||||||
sql += "ByLikes = MAX(IFNULL(ByLikes, 0), @byLikes) ";
|
sql += "ByLikes = MAX(IFNULL(ByLikes, 0), @byLikes) ";
|
||||||
sql += " WHERE BlogName = @BlogName AND PostID = @PostID AND (";
|
sql += " WHERE BlogName = @BlogName AND PostID = @PostID AND (";
|
||||||
sql += "(@postDate <> '.' AND IFNULL(postDate, '') <> @postDate) OR ";
|
sql += "(@postDate <> '.' AND IFNULL(postDate, '') <> @postDate) OR ";
|
||||||
@@ -1752,7 +1823,10 @@ namespace URLNotesGrabberCORE
|
|||||||
sql += "IFNULL(hasImage, 0) <> @hasImage OR ";
|
sql += "IFNULL(hasImage, 0) <> @hasImage OR ";
|
||||||
sql += "(@byLikes = 1 AND IFNULL(ByLikes, 0) = 0) OR ";
|
sql += "(@byLikes = 1 AND IFNULL(ByLikes, 0) = 0) OR ";
|
||||||
sql += "((@rootBlogName IS NOT NULL AND @rootBlogName <> '' AND @rootBlogName <> '.') AND IFNULL(RootBlogName, '') <> @rootBlogName) OR ";
|
sql += "((@rootBlogName IS NOT NULL AND @rootBlogName <> '' AND @rootBlogName <> '.') AND IFNULL(RootBlogName, '') <> @rootBlogName) OR ";
|
||||||
sql += "((@rootURL IS NOT NULL AND @rootURL <> '' AND @rootURL <> '.') AND IFNULL(RootURL, '') <> @rootURL)";
|
sql += "((@rootURL IS NOT NULL AND @rootURL <> '' AND @rootURL <> '.') AND IFNULL(RootURL, '') <> @rootURL) OR ";
|
||||||
|
// Without this the SET above is unreachable for a row whose content is already
|
||||||
|
// current: the UPDATE would not fire, and the type would stay NULL forever.
|
||||||
|
sql += "(PostType IS NULL AND @postType IS NOT NULL)";
|
||||||
sql += ")";
|
sql += ")";
|
||||||
|
|
||||||
using (SQLiteCommand command = new SQLiteCommand(sql, connection))
|
using (SQLiteCommand command = new SQLiteCommand(sql, connection))
|
||||||
@@ -1780,6 +1854,7 @@ namespace URLNotesGrabberCORE
|
|||||||
command.Parameters.AddWithValue("@rootURL", string.IsNullOrWhiteSpace(rootURL) ? (object)DBNull.Value : rootURL);
|
command.Parameters.AddWithValue("@rootURL", string.IsNullOrWhiteSpace(rootURL) ? (object)DBNull.Value : rootURL);
|
||||||
command.Parameters.AddWithValue("@hasImage", hasImage ? 1 : 0);
|
command.Parameters.AddWithValue("@hasImage", hasImage ? 1 : 0);
|
||||||
command.Parameters.AddWithValue("@byLikes", byLikes ? 1 : 0);
|
command.Parameters.AddWithValue("@byLikes", byLikes ? 1 : 0);
|
||||||
|
command.Parameters.AddWithValue("@postType", (object?)postType ?? DBNull.Value);
|
||||||
command.Parameters.AddWithValue("@BlogName", blogName);
|
command.Parameters.AddWithValue("@BlogName", blogName);
|
||||||
command.Parameters.AddWithValue("@PostID", postID);
|
command.Parameters.AddWithValue("@PostID", postID);
|
||||||
|
|
||||||
@@ -1959,7 +2034,7 @@ namespace URLNotesGrabberCORE
|
|||||||
// Only fan out to rows that match the SELECT criteria in GetRepliesWithFilledText (NULL/empty/legacy-'.'). Never overwrite '?' (confirmed-empty) or already-fetched text.
|
// Only fan out to rows that match the SELECT criteria in GetRepliesWithFilledText (NULL/empty/legacy-'.'). Never overwrite '?' (confirmed-empty) or already-fetched text.
|
||||||
// The ABS() term cannot use an index on TimeStamp, before or after the integer schema; the NoteBlogId probe is what keeps this off a full scan.
|
// The ABS() term cannot use an index on TimeStamp, before or after the integer schema; the NoteBlogId probe is what keeps this off a full scan.
|
||||||
string sql = "UPDATE Notes SET replyText = @replyText, DateModified = @dateModified " +
|
string sql = "UPDATE Notes SET replyText = @replyText, DateModified = @dateModified " +
|
||||||
"WHERE NoteBlogId = (SELECT BlogId FROM BlogNames WHERE BlogName = @noteBlogName) " +
|
"WHERE NoteBlogId = (SELECT BlogId FROM Blogs WHERE BlogName = @noteBlogName) " +
|
||||||
"AND ABS(TimeStamp - @TimeStamp) <= 5 " +
|
"AND ABS(TimeStamp - @TimeStamp) <= 5 " +
|
||||||
"AND TypeId = (SELECT TypeId FROM NoteTypes WHERE Type = 'reply') " +
|
"AND TypeId = (SELECT TypeId FROM NoteTypes WHERE Type = 'reply') " +
|
||||||
"AND (replyText IS NULL OR replyText = '' OR replyText = '.') " +
|
"AND (replyText IS NULL OR replyText = '' OR replyText = '.') " +
|
||||||
@@ -2005,7 +2080,7 @@ namespace URLNotesGrabberCORE
|
|||||||
connection.Open();
|
connection.Open();
|
||||||
|
|
||||||
string sql = "UPDATE Notes SET replyText = @replyText, DateModified = @dateModified " +
|
string sql = "UPDATE Notes SET replyText = @replyText, DateModified = @dateModified " +
|
||||||
"WHERE RootBlogId = (SELECT BlogId FROM BlogNames WHERE BlogName = @rootBlogName) " +
|
"WHERE RootBlogId = (SELECT BlogId FROM Blogs WHERE BlogName = @rootBlogName) " +
|
||||||
"AND PostID = @PostID " +
|
"AND PostID = @PostID " +
|
||||||
"AND TypeId = (SELECT TypeId FROM NoteTypes WHERE Type = 'reply') " +
|
"AND TypeId = (SELECT TypeId FROM NoteTypes WHERE Type = 'reply') " +
|
||||||
"AND IFNULL(replyText, '.') <> @replyText";
|
"AND IFNULL(replyText, '.') <> @replyText";
|
||||||
@@ -2089,6 +2164,8 @@ namespace URLNotesGrabberCORE
|
|||||||
cmd.ExecuteNonQuery();
|
cmd.ExecuteNonQuery();
|
||||||
Console.WriteLine("[Migration] Added PostType column to Posts table");
|
Console.WriteLine("[Migration] Added PostType column to Posts table");
|
||||||
}
|
}
|
||||||
|
|
||||||
|
BackfillMissingPostTypes(connection);
|
||||||
}
|
}
|
||||||
catch (Exception ex)
|
catch (Exception ex)
|
||||||
{
|
{
|
||||||
@@ -2096,6 +2173,65 @@ namespace URLNotesGrabberCORE
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// <summary>
|
||||||
|
/// Types rows that carry no PostType, inferring it from which content columns they hold.
|
||||||
|
///
|
||||||
|
/// These are posts harvested from notes and likes rather than read out of a TumblThree
|
||||||
|
/// export, so no filename ever described them and --ingest can never reach them: it only
|
||||||
|
/// types a post it meets inside a real .txt. Content is the only signal they have.
|
||||||
|
///
|
||||||
|
/// Runs on every migration pass and is idempotent -- it only touches PostType IS NULL,
|
||||||
|
/// so a row typed once is never revisited. Rows whose columns give no signal at all stay
|
||||||
|
/// NULL and are skipped by OutputMode.
|
||||||
|
///
|
||||||
|
/// Mirrors PostTypes.InferFromContent; the two must agree. Notably HasImage is not
|
||||||
|
/// consulted, because most text posts carry it.
|
||||||
|
/// </summary>
|
||||||
|
private static void BackfillMissingPostTypes(SQLiteConnection connection)
|
||||||
|
{
|
||||||
|
const string set = @"
|
||||||
|
UPDATE Posts SET PostType = CASE
|
||||||
|
WHEN Has(Question) AND Has(Answer) THEN 'answers'
|
||||||
|
WHEN Has(Quote) THEN 'quotes'
|
||||||
|
WHEN Has(Link) THEN 'links'
|
||||||
|
WHEN Has(AudioCaption) THEN 'audios'
|
||||||
|
WHEN Has(Body) THEN 'texts'
|
||||||
|
WHEN Has(PhotoURL) OR Has(PhotoCaption) THEN 'images'
|
||||||
|
ELSE NULL END
|
||||||
|
WHERE PostType IS NULL";
|
||||||
|
|
||||||
|
// SQLite has no user-defined predicate here, so expand the "field supplied" test
|
||||||
|
// ("." is the not-supplied sentinel used throughout the export format) inline.
|
||||||
|
string sql = System.Text.RegularExpressions.Regex.Replace(
|
||||||
|
set, @"Has\((\w+)\)", "TRIM(IFNULL($1, '')) NOT IN ('', '.')");
|
||||||
|
|
||||||
|
try
|
||||||
|
{
|
||||||
|
long before;
|
||||||
|
using (var count = new SQLiteCommand("SELECT COUNT(*) FROM Posts WHERE PostType IS NULL", connection))
|
||||||
|
before = Convert.ToInt64(count.ExecuteScalar());
|
||||||
|
|
||||||
|
if (before == 0) return;
|
||||||
|
|
||||||
|
int changed;
|
||||||
|
using (var cmd = new SQLiteCommand(sql, connection))
|
||||||
|
changed = cmd.ExecuteNonQuery();
|
||||||
|
|
||||||
|
long after;
|
||||||
|
using (var count = new SQLiteCommand("SELECT COUNT(*) FROM Posts WHERE PostType IS NULL", connection))
|
||||||
|
after = Convert.ToInt64(count.ExecuteScalar());
|
||||||
|
|
||||||
|
if (changed > 0 || after != before)
|
||||||
|
Console.WriteLine($"[Migration] Backfilled PostType for {before - after} post(s); {after} still untyped (no content signal).");
|
||||||
|
}
|
||||||
|
catch (Exception ex)
|
||||||
|
{
|
||||||
|
// A failed backfill must not stop the run: untyped rows are skipped on export,
|
||||||
|
// which is inconvenient, not corrupting.
|
||||||
|
Console.WriteLine($"[Migration] PostType backfill failed: {ex.Message}");
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
// INSERT-or-UPDATE for a post arriving from a Tumblr text-file export.
|
// INSERT-or-UPDATE for a post arriving from a Tumblr text-file export.
|
||||||
// On collision, only content columns + PostType + DateModified are updated;
|
// On collision, only content columns + PostType + DateModified are updated;
|
||||||
// engagement columns (ByLikes, RootBlogName, RootURL, HasNotesGathered, NotFound,
|
// engagement columns (ByLikes, RootBlogName, RootURL, HasNotesGathered, NotFound,
|
||||||
@@ -2126,6 +2262,11 @@ namespace URLNotesGrabberCORE
|
|||||||
string? DBPath = null)
|
string? DBPath = null)
|
||||||
{
|
{
|
||||||
DBPath ??= GetDefaultDbPath();
|
DBPath ??= GetDefaultDbPath();
|
||||||
|
// Central guarantee: whatever a caller believes, only a canonical type reaches the
|
||||||
|
// column. PostType is used as an output filename, so this is the invariant that keeps
|
||||||
|
// a stray value from becoming a stray file.
|
||||||
|
postType = PostTypes.Normalize(postType);
|
||||||
|
postDate = PostDates.Normalize(postDate);
|
||||||
try { AddBlog(blogName, false, DBPath); } catch { }
|
try { AddBlog(blogName, false, DBPath); } catch { }
|
||||||
|
|
||||||
SQLiteConnection connection;
|
SQLiteConnection connection;
|
||||||
@@ -2497,10 +2638,11 @@ namespace URLNotesGrabberCORE
|
|||||||
if (string.IsNullOrWhiteSpace(kvp.Value)) continue;
|
if (string.IsNullOrWhiteSpace(kvp.Value)) continue;
|
||||||
string? column = MapPrefixToColumn(kvp.Key);
|
string? column = MapPrefixToColumn(kvp.Key);
|
||||||
if (column == null) continue;
|
if (column == null) continue;
|
||||||
|
string value = column == "PostDate" ? PostDates.Normalize(kvp.Value)! : kvp.Value;
|
||||||
string paramName = "@p" + parameters.Count;
|
string paramName = "@p" + parameters.Count;
|
||||||
setClauses.Add($"{column} = {paramName}");
|
setClauses.Add($"{column} = {paramName}");
|
||||||
changedClauses.Add($"IFNULL({column}, '') <> {paramName}");
|
changedClauses.Add($"IFNULL({column}, '') <> {paramName}");
|
||||||
parameters.Add((paramName, kvp.Value));
|
parameters.Add((paramName, value));
|
||||||
}
|
}
|
||||||
|
|
||||||
if (setClauses.Count == 0) return false;
|
if (setClauses.Count == 0) return false;
|
||||||
|
|||||||
@@ -85,7 +85,25 @@ namespace URLNotesGrabberCORE
|
|||||||
{
|
{
|
||||||
string rawBlogName = Path.GetFileName(Path.GetDirectoryName(file) ?? "unknown");
|
string rawBlogName = Path.GetFileName(Path.GetDirectoryName(file) ?? "unknown");
|
||||||
string blogName = Regex.Replace(rawBlogName, @"_\d+$", "");
|
string blogName = Regex.Replace(rawBlogName, @"_\d+$", "");
|
||||||
string postType = Path.GetFileNameWithoutExtension(file);
|
|
||||||
|
// The filename becomes the row's PostType, and PostType later becomes an
|
||||||
|
// output filename -- so an unrecognized name here would mint a new type and
|
||||||
|
// a new file from any stray .txt that happens to sit in the tree. Only the
|
||||||
|
// eight real export files are ingestable.
|
||||||
|
//
|
||||||
|
// This is also what breaks the Unknown.txt cycle: OutputMode used to write
|
||||||
|
// untyped rows to Unknown.txt, and this scan would read it straight back
|
||||||
|
// and stamp those rows with the literal type "Unknown", making the file
|
||||||
|
// regenerate itself forever.
|
||||||
|
string? resolvedPostType = PostTypes.FromFileName(file);
|
||||||
|
if (resolvedPostType == null)
|
||||||
|
{
|
||||||
|
filesSkipped++;
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
// Non-nullable from here so the local Flush() below stays warning-clean:
|
||||||
|
// nullable flow analysis does not reach into local functions.
|
||||||
|
string postType = resolvedPostType;
|
||||||
|
|
||||||
if (targetBlog != null && !string.Equals(blogName, targetBlog, StringComparison.OrdinalIgnoreCase))
|
if (targetBlog != null && !string.Equals(blogName, targetBlog, StringComparison.OrdinalIgnoreCase))
|
||||||
{
|
{
|
||||||
|
|||||||
@@ -117,7 +117,12 @@ namespace URLNotesGrabberCORE
|
|||||||
question: reader.IsDBNull(18) ? null : reader.GetString(18),
|
question: reader.IsDBNull(18) ? null : reader.GetString(18),
|
||||||
answer: reader.IsDBNull(19) ? null : reader.GetString(19),
|
answer: reader.IsDBNull(19) ? null : reader.GetString(19),
|
||||||
title: reader.IsDBNull(20) ? null : reader.GetString(20),
|
title: reader.IsDBNull(20) ? null : reader.GetString(20),
|
||||||
postType: reader.IsDBNull(21) ? null : reader.GetString(21),
|
// A legacy Posts.db predating the PostType column hands back NULL
|
||||||
|
// here, and on the INSERT branch that NULL is stored -- reseeding
|
||||||
|
// exactly the untyped rows the backfill exists to clear. Normalize
|
||||||
|
// so an unrecognized legacy value cannot become a filename either;
|
||||||
|
// the backfill types whatever comes through as null.
|
||||||
|
postType: PostTypes.Normalize(reader.IsDBNull(21) ? null : reader.GetString(21)),
|
||||||
hasImage: hasImage);
|
hasImage: hasImage);
|
||||||
postsUpserted++;
|
postsUpserted++;
|
||||||
if (postsUpserted % 500 == 0)
|
if (postsUpserted % 500 == 0)
|
||||||
|
|||||||
@@ -65,10 +65,20 @@ namespace URLNotesGrabberCORE
|
|||||||
var posts = DataAccess.GetAllPostsForBlog(blogName);
|
var posts = DataAccess.GetAllPostsForBlog(blogName);
|
||||||
Console.WriteLine($" Found {posts.Count} post(s) for this blog.");
|
Console.WriteLine($" Found {posts.Count} post(s) for this blog.");
|
||||||
|
|
||||||
var grouped = posts.GroupBy(p => p.PostType ?? "Unknown");
|
// A post's type becomes a filename, so only a recognized type may be written. The
|
||||||
|
// old `PostType ?? "Unknown"` invented Unknown.txt for untyped rows, which --ingest
|
||||||
|
// then read back as a type named "Unknown" -- the two regenerated each other.
|
||||||
|
// Untyped rows are skipped instead: after the backfill these are only rows with no
|
||||||
|
// content signal at all, so nothing meaningful is lost, and nothing is invented.
|
||||||
|
var typed = posts.Where(p => PostTypes.Normalize(p.PostType) != null).ToList();
|
||||||
|
int untyped = posts.Count - typed.Count;
|
||||||
|
if (untyped > 0)
|
||||||
|
Console.WriteLine($" Skipping {untyped} post(s) with no recognized PostType.");
|
||||||
|
|
||||||
|
var grouped = typed.GroupBy(p => PostTypes.Normalize(p.PostType)!);
|
||||||
foreach (var typeGroup in grouped)
|
foreach (var typeGroup in grouped)
|
||||||
{
|
{
|
||||||
string postType = typeGroup.Key ?? "Unknown";
|
string postType = typeGroup.Key;
|
||||||
string outputFilePath = Path.Combine(folder, $"{postType}.txt");
|
string outputFilePath = Path.Combine(folder, $"{postType}.txt");
|
||||||
var ordered = typeGroup.OrderBy(p => p.Date).ToList();
|
var ordered = typeGroup.OrderBy(p => p.Date).ToList();
|
||||||
Console.WriteLine($" Writing {ordered.Count} post(s) to {postType}.txt");
|
Console.WriteLine($" Writing {ordered.Count} post(s) to {postType}.txt");
|
||||||
|
|||||||
@@ -0,0 +1,27 @@
|
|||||||
|
using System;
|
||||||
|
using System.Globalization;
|
||||||
|
|
||||||
|
namespace URLNotesGrabberCORE
|
||||||
|
{
|
||||||
|
/// <summary>
|
||||||
|
/// The single format for Posts.PostDate: "yyyy-MM-dd HH:mm:ss GMT", which is what the
|
||||||
|
/// Tumblr API sends and what nearly every row holds. Text-file exports can carry the
|
||||||
|
/// RFC 1123 form ("Fri, 14 Feb 2025 15:20:09 GMT") instead, which as text sorts on its
|
||||||
|
/// weekday name and falls outside every --fromDate/--toDate range comparison.
|
||||||
|
///
|
||||||
|
/// Every path that writes PostDate routes through <see cref="Normalize"/>. Only the RFC 1123
|
||||||
|
/// form is rewritten; anything else, including the "." no-change sentinel, passes through.
|
||||||
|
/// </summary>
|
||||||
|
public static class PostDates
|
||||||
|
{
|
||||||
|
public static string? Normalize(string? value)
|
||||||
|
{
|
||||||
|
if (string.IsNullOrWhiteSpace(value)) return value;
|
||||||
|
string trimmed = value.Trim();
|
||||||
|
if (DateTime.TryParseExact(trimmed, "r", CultureInfo.InvariantCulture,
|
||||||
|
DateTimeStyles.AdjustToUniversal | DateTimeStyles.AssumeUniversal, out DateTime parsed))
|
||||||
|
return parsed.ToString("yyyy-MM-dd HH:mm:ss", CultureInfo.InvariantCulture) + " GMT";
|
||||||
|
return value;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,123 @@
|
|||||||
|
using System;
|
||||||
|
using System.Collections.Generic;
|
||||||
|
using System.IO;
|
||||||
|
|
||||||
|
namespace URLNotesGrabberCORE
|
||||||
|
{
|
||||||
|
/// <summary>
|
||||||
|
/// The single source of truth for Posts.PostType values.
|
||||||
|
///
|
||||||
|
/// PostType exists so --output can write one .txt per type. Because the type becomes a
|
||||||
|
/// *filename*, an unvalidated value is not a cosmetic problem: it creates a file. That is
|
||||||
|
/// how "Unknown.txt" came about -- OutputMode used `PostType ?? "Unknown"` as a filename,
|
||||||
|
/// --ingest then read that file straight back and derived the literal type "Unknown" from
|
||||||
|
/// its name, and the pair would have kept regenerating each other indefinitely.
|
||||||
|
///
|
||||||
|
/// So every path that produces a type routes through <see cref="Normalize"/>, which admits
|
||||||
|
/// only the eight known names and returns null for anything else. A null type is safe:
|
||||||
|
/// OutputMode skips those rows rather than inventing a file for them.
|
||||||
|
/// </summary>
|
||||||
|
public static class PostTypes
|
||||||
|
{
|
||||||
|
// The canonical set. These are exactly the TumblThree .txt basenames, which is what
|
||||||
|
// makes an ingested filename usable as a type without translation.
|
||||||
|
public const string Texts = "texts";
|
||||||
|
public const string Answers = "answers";
|
||||||
|
public const string Quotes = "quotes";
|
||||||
|
public const string Links = "links";
|
||||||
|
public const string Conversations = "conversations";
|
||||||
|
public const string Images = "images";
|
||||||
|
public const string Videos = "videos";
|
||||||
|
public const string Audios = "audios";
|
||||||
|
|
||||||
|
private static readonly HashSet<string> Known = new HashSet<string>(
|
||||||
|
new[] { Texts, Answers, Quotes, Links, Conversations, Images, Videos, Audios },
|
||||||
|
StringComparer.OrdinalIgnoreCase);
|
||||||
|
|
||||||
|
// Tumblr's legacy post format (npf=false) names types in the singular. The likes API is
|
||||||
|
// the one source that reports a type directly rather than via a filename, so it is the
|
||||||
|
// only place this mapping is needed.
|
||||||
|
private static readonly Dictionary<string, string> ApiTypeMap = new Dictionary<string, string>(StringComparer.OrdinalIgnoreCase)
|
||||||
|
{
|
||||||
|
["text"] = Texts,
|
||||||
|
["photo"] = Images,
|
||||||
|
["quote"] = Quotes,
|
||||||
|
["link"] = Links,
|
||||||
|
["chat"] = Conversations,
|
||||||
|
["answer"] = Answers,
|
||||||
|
["audio"] = Audios,
|
||||||
|
["video"] = Videos,
|
||||||
|
};
|
||||||
|
|
||||||
|
/// <summary>
|
||||||
|
/// Returns the canonical type name, or null if the value is not one of the eight.
|
||||||
|
/// Returning null rather than passing the value through is the whole point: an
|
||||||
|
/// unrecognized string must never reach a filename.
|
||||||
|
/// </summary>
|
||||||
|
public static string? Normalize(string? candidate)
|
||||||
|
{
|
||||||
|
if (string.IsNullOrWhiteSpace(candidate)) return null;
|
||||||
|
string trimmed = candidate.Trim();
|
||||||
|
return Known.TryGetValue(trimmed, out string? canonical) ? canonical : null;
|
||||||
|
}
|
||||||
|
|
||||||
|
/// <summary>
|
||||||
|
/// Type for a post read out of a TumblThree export file, taken from the filename
|
||||||
|
/// ("texts.txt" -> "texts"). Anything else in the folder -- README.txt, a stray
|
||||||
|
/// triage file, or a previously written Unknown.txt -- normalizes to null and is
|
||||||
|
/// rejected by the caller.
|
||||||
|
/// </summary>
|
||||||
|
public static string? FromFileName(string? path)
|
||||||
|
{
|
||||||
|
if (string.IsNullOrWhiteSpace(path)) return null;
|
||||||
|
return Normalize(Path.GetFileNameWithoutExtension(path));
|
||||||
|
}
|
||||||
|
|
||||||
|
/// <summary>
|
||||||
|
/// Type for a post from the likes API, whose legacy-format `type` field is singular.
|
||||||
|
/// Null when the field is absent or unrecognized -- the access is dynamic, so a missing
|
||||||
|
/// field yields null at runtime rather than failing to compile.
|
||||||
|
/// </summary>
|
||||||
|
public static string? FromApiType(string? apiType)
|
||||||
|
{
|
||||||
|
if (string.IsNullOrWhiteSpace(apiType)) return null;
|
||||||
|
return ApiTypeMap.TryGetValue(apiType.Trim(), out string? mapped) ? mapped : null;
|
||||||
|
}
|
||||||
|
|
||||||
|
/// <summary>
|
||||||
|
/// Last-resort type inferred from which content columns a row actually carries. Used
|
||||||
|
/// only to backfill rows written before any type was recorded; a filename or an API
|
||||||
|
/// type is always preferred over this.
|
||||||
|
///
|
||||||
|
/// The order matters and is derived from the already-typed rows, where the column
|
||||||
|
/// signatures are effectively disjoint: answers carry Question+Answer and no Body,
|
||||||
|
/// images carry photo columns and no Body, texts carry Body and no photo columns.
|
||||||
|
///
|
||||||
|
/// HasImage is deliberately NOT consulted: it is set on 12,420 of 19,828 known text
|
||||||
|
/// posts, so it says nothing about the post's type.
|
||||||
|
///
|
||||||
|
/// conversations cannot be separated from texts this way -- both carry only Body -- so
|
||||||
|
/// a chat post with no other signal is labelled texts. A later --ingest that meets the
|
||||||
|
/// post in a real conversations.txt corrects it.
|
||||||
|
/// </summary>
|
||||||
|
public static string? InferFromContent(string? question, string? answer, string? quote,
|
||||||
|
string? link, string? audioCaption, string? body, string? photoUrl, string? photoCaption)
|
||||||
|
{
|
||||||
|
if (HasValue(question) && HasValue(answer)) return Answers;
|
||||||
|
if (HasValue(quote)) return Quotes;
|
||||||
|
if (HasValue(link)) return Links;
|
||||||
|
if (HasValue(audioCaption)) return Audios;
|
||||||
|
if (HasValue(body)) return Texts;
|
||||||
|
if (HasValue(photoUrl) || HasValue(photoCaption)) return Images;
|
||||||
|
return null;
|
||||||
|
}
|
||||||
|
|
||||||
|
// "." is the codebase-wide "field not supplied" sentinel in export records, so it
|
||||||
|
// counts as absent here just as it does in UpdatePost's CASE guards.
|
||||||
|
private static bool HasValue(string? value)
|
||||||
|
{
|
||||||
|
if (string.IsNullOrWhiteSpace(value)) return false;
|
||||||
|
return value.Trim() != ".";
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -44,6 +44,8 @@ namespace URLNotesGrabberCORE
|
|||||||
bool apiExplicitlySet = false;
|
bool apiExplicitlySet = false;
|
||||||
string startFromBlogName = string.Empty;
|
string startFromBlogName = string.Empty;
|
||||||
bool forceIgnoreCooldown = false;
|
bool forceIgnoreCooldown = false;
|
||||||
|
DateTime? fromDate = null;
|
||||||
|
DateTime? toDate = null;
|
||||||
List<string> filteredArgs = new List<string>();
|
List<string> filteredArgs = new List<string>();
|
||||||
for (int i = 0; i < args.Length; i++)
|
for (int i = 0; i < args.Length; i++)
|
||||||
{
|
{
|
||||||
@@ -60,6 +62,34 @@ namespace URLNotesGrabberCORE
|
|||||||
continue;
|
continue;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
if (string.Equals(args[i], "--fromDate", StringComparison.OrdinalIgnoreCase))
|
||||||
|
{
|
||||||
|
if (i + 1 < args.Length && DateTime.TryParse(args[i + 1], out DateTime parsedFromDate))
|
||||||
|
{
|
||||||
|
fromDate = parsedFromDate;
|
||||||
|
i++;
|
||||||
|
}
|
||||||
|
else
|
||||||
|
{
|
||||||
|
Console.WriteLine("--Missing or unparseable date after --fromDate. Ignoring.--");
|
||||||
|
}
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
|
||||||
|
if (string.Equals(args[i], "--toDate", StringComparison.OrdinalIgnoreCase))
|
||||||
|
{
|
||||||
|
if (i + 1 < args.Length && DateTime.TryParse(args[i + 1], out DateTime parsedToDate))
|
||||||
|
{
|
||||||
|
toDate = parsedToDate;
|
||||||
|
i++;
|
||||||
|
}
|
||||||
|
else
|
||||||
|
{
|
||||||
|
Console.WriteLine("--Missing or unparseable date after --toDate. Ignoring.--");
|
||||||
|
}
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
|
||||||
if (string.Equals(args[i], "--api3", StringComparison.OrdinalIgnoreCase))
|
if (string.Equals(args[i], "--api3", StringComparison.OrdinalIgnoreCase))
|
||||||
{
|
{
|
||||||
apiSectionName = "TumblrApi3";
|
apiSectionName = "TumblrApi3";
|
||||||
@@ -148,6 +178,12 @@ namespace URLNotesGrabberCORE
|
|||||||
if (args.Length == 0) //Traverse folder structure to add posts and thus blogs to DB
|
if (args.Length == 0) //Traverse folder structure to add posts and thus blogs to DB
|
||||||
{
|
{
|
||||||
int postsAdded = 0;
|
int postsAdded = 0;
|
||||||
|
// This is the mode that actually gets run day to day, so the schema migration and
|
||||||
|
// the PostType backfill have to happen here too. They used to hang off --ingest,
|
||||||
|
// --output and friends only, which meant the untyped rows this traversal creates
|
||||||
|
// could sit unrepaired indefinitely while the one command everyone runs skipped
|
||||||
|
// the fix entirely. Idempotent, so paying it on every run costs nothing.
|
||||||
|
DataAccess.EnsureTTFileHelperColumnsExist();
|
||||||
try
|
try
|
||||||
{
|
{
|
||||||
DataAccess.EnableImportModePragmas();
|
DataAccess.EnableImportModePragmas();
|
||||||
@@ -316,7 +352,22 @@ namespace URLNotesGrabberCORE
|
|||||||
managedCollectRun = true;
|
managedCollectRun = true;
|
||||||
}
|
}
|
||||||
|
|
||||||
exitCode = CollectNotes(settings.GetValue<string>("PathOutput"), withoutNotesOnly, beforeDate, managedCollectRun, collectBlogName).GetAwaiter().GetResult();
|
if (forceIgnoreCooldown)
|
||||||
|
Console.WriteLine(withoutNotesOnly
|
||||||
|
? "--force: ignoring the periodic re-collect cooldown; already-collected posts in scope are re-queued now"
|
||||||
|
: "--force: no effect in mode 0 - a full re-check already re-collects every post");
|
||||||
|
|
||||||
|
if (fromDate.HasValue)
|
||||||
|
Console.WriteLine(withoutNotesOnly
|
||||||
|
? $"--fromDate: only re-queuing already-collected posts originally posted on/after {fromDate.Value} (applies with or without --force)"
|
||||||
|
: "--fromDate: no effect in mode 0 - it only bounds the periodic re-queue branch");
|
||||||
|
|
||||||
|
if (toDate.HasValue)
|
||||||
|
Console.WriteLine(withoutNotesOnly
|
||||||
|
? $"--toDate: only re-queuing already-collected posts originally posted on/before {toDate.Value} (applies with or without --force)"
|
||||||
|
: "--toDate: no effect in mode 0 - it only bounds the periodic re-queue branch");
|
||||||
|
|
||||||
|
exitCode = CollectNotes(settings.GetValue<string>("PathOutput"), withoutNotesOnly, beforeDate, managedCollectRun, collectBlogName, forceIgnoreCooldown, fromDate, toDate).GetAwaiter().GetResult();
|
||||||
break;
|
break;
|
||||||
|
|
||||||
case "--blogsR": //collect notes from all posts
|
case "--blogsR": //collect notes from all posts
|
||||||
@@ -448,7 +499,7 @@ namespace URLNotesGrabberCORE
|
|||||||
|
|
||||||
Console.WriteLine("--blogs\t For each Blog in DB, write blogname to file");
|
Console.WriteLine("--blogs\t For each Blog in DB, write blogname to file");
|
||||||
|
|
||||||
Console.WriteLine("--collect [0|1] [datetime] [blogname]\t Collect Notes from API. 1=only posts without notes. 0=full re-check of all posts: a single resumable pass (interrupt & relaunch to resume; stops when complete, retrigger for a new pass). Optional datetime overrides the cutoff and runs as a one-off (bypasses resume tracking). Optional blogname restricts the run to that blog (exact, case-sensitive match) and also runs as a one-off; e.g. \"--collect 1 zomb-eh\". datetime and blogname may be given in either order - use --blog=name if a blog name would otherwise parse as a date.");
|
Console.WriteLine("--collect [0|1] [datetime] [blogname]\t Collect Notes from API. 1=only posts without notes. 0=full re-check of all posts: a single resumable pass (interrupt & relaunch to resume; stops when complete, retrigger for a new pass). Optional datetime overrides the cutoff and runs as a one-off (bypasses resume tracking). Optional blogname restricts the run to that blog (exact, case-sensitive match) and also runs as a one-off; e.g. \"--collect 1 zomb-eh\". datetime and blogname may be given in either order - use --blog=name if a blog name would otherwise parse as a date. Add --force to ignore the periodic re-collect cooldown and re-queue already-collected posts immediately (mode 1 only). Add --fromDate <datetime> / --toDate <datetime> to only re-queue already-collected posts originally posted on/after / on/before that date (mode 1 only; either or both may be given; applies with or without --force).");
|
||||||
|
|
||||||
Console.WriteLine("--blogsR\t For each Note that is a REPLY, write blogname to file ");
|
Console.WriteLine("--blogsR\t For each Note that is a REPLY, write blogname to file ");
|
||||||
|
|
||||||
@@ -460,7 +511,11 @@ namespace URLNotesGrabberCORE
|
|||||||
|
|
||||||
Console.WriteLine("--likes\t Fetch likes: initial backfill for new blogs, incremental refresh for blogs past cooldown. Optional blog name forces single-blog run.");
|
Console.WriteLine("--likes\t Fetch likes: initial backfill for new blogs, incremental refresh for blogs past cooldown. Optional blog name forces single-blog run.");
|
||||||
|
|
||||||
Console.WriteLine("--force\t (with --likes) Ignore cooldown and refresh every fully-backfilled blog");
|
Console.WriteLine("--force\t Ignore refresh cooldowns: with --likes, refresh every fully-backfilled blog; with --collect 1, re-queue already-collected posts without waiting out their cooldown");
|
||||||
|
|
||||||
|
Console.WriteLine("--fromDate <datetime>\t With --collect 1, only re-queue already-collected posts originally posted on/after <datetime>. Independent of --force - applies whether or not the cooldown is also bypassed.");
|
||||||
|
|
||||||
|
Console.WriteLine("--toDate <datetime>\t With --collect 1, only re-queue already-collected posts originally posted on/before <datetime>. Independent of --force; may be combined with --fromDate for a range.");
|
||||||
|
|
||||||
Console.WriteLine("--urldump\t Scan all posts' text columns and extract suspected URLs to configured file");
|
Console.WriteLine("--urldump\t Scan all posts' text columns and extract suspected URLs to configured file");
|
||||||
|
|
||||||
@@ -1047,6 +1102,15 @@ namespace URLNotesGrabberCORE
|
|||||||
string reblogKey = post.reblog_key?.ToString() ?? ".";
|
string reblogKey = post.reblog_key?.ToString() ?? ".";
|
||||||
string link = ".";
|
string link = ".";
|
||||||
|
|
||||||
|
// No file backs a liked post, so the filename trick used everywhere
|
||||||
|
// else cannot apply here. GrabLikes requests npf=false, and in the
|
||||||
|
// legacy format `type` is the discriminator that decides which content
|
||||||
|
// fields a post carries -- singular there, mapped to our plural names.
|
||||||
|
// liked_posts is List<dynamic>, so this is resolved at runtime and a
|
||||||
|
// missing field yields null rather than a compile error; an absent or
|
||||||
|
// unrecognized value leaves the type NULL instead of guessing.
|
||||||
|
string? apiPostType = PostTypes.FromApiType(post.type?.ToString() as string);
|
||||||
|
|
||||||
// Only insert if any of the data contains strings from ContainsList
|
// Only insert if any of the data contains strings from ContainsList
|
||||||
bool shouldInsert = false;
|
bool shouldInsert = false;
|
||||||
string matchedFieldName = string.Empty;
|
string matchedFieldName = string.Empty;
|
||||||
@@ -1090,7 +1154,8 @@ if (shouldInsert)
|
|||||||
DataAccess.AddPost(authorBlog, postID, reblogURL, date, postURL, slug, reblogKey,
|
DataAccess.AddPost(authorBlog, postID, reblogURL, date, postURL, slug, reblogKey,
|
||||||
reblogName, summary, quote, body, tags, link, photoURL,
|
reblogName, summary, quote, body, tags, link, photoURL,
|
||||||
photoCaption, downloadedFiles, audioCaption, question, answer,
|
photoCaption, downloadedFiles, audioCaption, question, answer,
|
||||||
title, hasImage, true, rootBlogName: rootBlogName, rootURL: rootURL);
|
title, hasImage, true, rootBlogName: rootBlogName, rootURL: rootURL,
|
||||||
|
postType: apiPostType);
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -1329,15 +1394,17 @@ if (shouldInsert)
|
|||||||
// blipping on one post. Past this, skipping post-by-post would just hammer a closed door.
|
// blipping on one post. Past this, skipping post-by-post would just hammer a closed door.
|
||||||
const int MaxConsecutiveTransient = 10;
|
const int MaxConsecutiveTransient = 10;
|
||||||
|
|
||||||
static async Task<int> CollectNotes(string outPath, bool withoutNotesOnly = true, DateTime? beforeDate = null, bool managedRun = false, string? blogName = null)
|
static async Task<int> CollectNotes(string outPath, bool withoutNotesOnly = true, DateTime? beforeDate = null, bool managedRun = false, string? blogName = null, bool ignoreRefreshCooldown = false, DateTime? fromDate = null, DateTime? toDate = null)
|
||||||
{
|
{
|
||||||
List<Tuple<string, long, long, long>> posts = DataAccess.GetPosts(withoutNotesOnly, beforeDate, blogName);
|
List<Tuple<string, long, long, long>> posts = DataAccess.GetPosts(withoutNotesOnly, beforeDate, blogName, ignoreRefreshCooldown, fromDate, toDate);
|
||||||
|
|
||||||
if (posts.Count == 0 && !string.IsNullOrWhiteSpace(blogName))
|
if (posts.Count == 0 && !string.IsNullOrWhiteSpace(blogName))
|
||||||
{
|
{
|
||||||
// BlogName is matched exactly, so a typo or a case mismatch looks identical to "nothing
|
// BlogName is matched exactly, so a typo or a case mismatch looks identical to "nothing
|
||||||
// left to collect". Say so rather than reporting a silent, instant success.
|
// left to collect". Say so rather than reporting a silent, instant success.
|
||||||
Console.WriteLine($"No posts to collect for blog '{blogName}'. Either it is fully collected, or the name does not match a stored blog (the match is case-sensitive).");
|
Console.WriteLine($"No posts to collect for blog '{blogName}'. Either it is fully collected, or the name does not match a stored blog (the match is case-sensitive).");
|
||||||
|
if (withoutNotesOnly && !ignoreRefreshCooldown)
|
||||||
|
Console.WriteLine("Already-collected posts are re-queued only once their cooldown elapses; add --force to re-collect them now.");
|
||||||
return 0;
|
return 0;
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -1435,7 +1502,7 @@ if (shouldInsert)
|
|||||||
}
|
}
|
||||||
|
|
||||||
// Re-fetch the updated list after processing the current post
|
// Re-fetch the updated list after processing the current post
|
||||||
posts = DataAccess.GetPosts(withoutNotesOnly, beforeDate, blogName);
|
posts = DataAccess.GetPosts(withoutNotesOnly, beforeDate, blogName, ignoreRefreshCooldown, fromDate, toDate);
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -1508,7 +1575,15 @@ if (shouldInsert)
|
|||||||
string normalizedDirectoryName = NormalizeBlogFolderName(new DirectoryInfo(path).Name);
|
string normalizedDirectoryName = NormalizeBlogFolderName(new DirectoryInfo(path).Name);
|
||||||
bool isAtOrAfterStart = string.IsNullOrWhiteSpace(startFromBlogName) || string.Compare(normalizedDirectoryName, startFromBlogName, StringComparison.OrdinalIgnoreCase) >= 0;
|
bool isAtOrAfterStart = string.IsNullOrWhiteSpace(startFromBlogName) || string.Compare(normalizedDirectoryName, startFromBlogName, StringComparison.OrdinalIgnoreCase) >= 0;
|
||||||
|
|
||||||
|
// The filename is the post type ("texts.txt" -> "texts"), so only the eight
|
||||||
|
// known export files are post sources. Everything else in a blog folder is
|
||||||
|
// either not a post file at all (README.txt, url lists, triage scratch) or
|
||||||
|
// is our own derived output -- Unknown.txt above all, which must never be
|
||||||
|
// read back in as a source or it perpetuates itself.
|
||||||
|
string? filePostType = PostTypes.FromFileName(file);
|
||||||
|
|
||||||
if (file.EndsWith(".txt", StringComparison.OrdinalIgnoreCase)
|
if (file.EndsWith(".txt", StringComparison.OrdinalIgnoreCase)
|
||||||
|
&& filePostType != null
|
||||||
&& (string.IsNullOrEmpty(blogName) || path.IndexOf(blogName, StringComparison.OrdinalIgnoreCase) >= 0)
|
&& (string.IsNullOrEmpty(blogName) || path.IndexOf(blogName, StringComparison.OrdinalIgnoreCase) >= 0)
|
||||||
&& isAtOrAfterStart)
|
&& isAtOrAfterStart)
|
||||||
{
|
{
|
||||||
@@ -1541,7 +1616,7 @@ if (shouldInsert)
|
|||||||
DataAccess.AddPost(curDir, long.Parse(reblog.postID), reblog.reblogURL, reblog.date, reblog.postURL, reblog.slug, reblog.reblogKey,
|
DataAccess.AddPost(curDir, long.Parse(reblog.postID), reblog.reblogURL, reblog.date, reblog.postURL, reblog.slug, reblog.reblogKey,
|
||||||
reblog.reblogName, reblog.summary, reblog.quote, reblog.body, reblog.tags, reblog.link, reblog.photoURL,
|
reblog.reblogName, reblog.summary, reblog.quote, reblog.body, reblog.tags, reblog.link, reblog.photoURL,
|
||||||
reblog.photoCaption, reblog.downloadedFiles, reblog.audioCaption, reblog.question, reblog.answer,
|
reblog.photoCaption, reblog.downloadedFiles, reblog.audioCaption, reblog.question, reblog.answer,
|
||||||
reblog.title, false, rootURL: reblog.rootURL);
|
reblog.title, false, rootURL: reblog.rootURL, postType: filePostType);
|
||||||
recordImportStopwatch.Stop();
|
recordImportStopwatch.Stop();
|
||||||
|
|
||||||
postsAdded++;
|
postsAdded++;
|
||||||
@@ -1689,7 +1764,7 @@ if (shouldInsert)
|
|||||||
DataAccess.AddPost(curDir, long.Parse(reblog.postID), reblog.reblogURL, reblog.date, reblog.postURL, reblog.slug, reblog.reblogKey,
|
DataAccess.AddPost(curDir, long.Parse(reblog.postID), reblog.reblogURL, reblog.date, reblog.postURL, reblog.slug, reblog.reblogKey,
|
||||||
reblog.reblogName, reblog.summary, reblog.quote, reblog.body, reblog.tags, reblog.link, reblog.photoURL,
|
reblog.reblogName, reblog.summary, reblog.quote, reblog.body, reblog.tags, reblog.link, reblog.photoURL,
|
||||||
reblog.photoCaption, reblog.downloadedFiles, reblog.audioCaption, reblog.question, reblog.answer,
|
reblog.photoCaption, reblog.downloadedFiles, reblog.audioCaption, reblog.question, reblog.answer,
|
||||||
reblog.title, true, rootURL: reblog.rootURL);
|
reblog.title, true, rootURL: reblog.rootURL, postType: filePostType);
|
||||||
recordImportStopwatch.Stop();
|
recordImportStopwatch.Stop();
|
||||||
|
|
||||||
postsAdded++;
|
postsAdded++;
|
||||||
|
|||||||
+116
-94
@@ -24,21 +24,43 @@ Everything below was read out of the live file, not inferred from code. Counts a
|
|||||||
> Applied by `../normalize-notes.sql`, which took the file from 207 MB to 148 MB. An
|
> Applied by `../normalize-notes.sql`, which took the file from 207 MB to 148 MB. An
|
||||||
> earlier change the same day (`../shrink-db.sql`) took it from 267 MB to 207 MB.
|
> earlier change the same day (`../shrink-db.sql`) took it from 267 MB to 207 MB.
|
||||||
|
|
||||||
|
> ### ⚠ Breaking change, 2026-09-28: `BlogNames` is gone; `Blogs.BlogId` is the only ID authority
|
||||||
|
>
|
||||||
|
> The IDs in `Notes` used to live in a `BlogNames` table, with a copy in `Blogs.BlogId`.
|
||||||
|
> Nothing kept the copy current, so by 2026-09-28 12,238 blogs first seen after the
|
||||||
|
> migration had `Blogs.BlogId = NULL`. Every `Notes`-to-`Blogs` join on `BlogId` silently
|
||||||
|
> skipped them and their 23,148 notes, which kept them out of `GetBlogs`.
|
||||||
|
>
|
||||||
|
> `../retire-blognames.sql` fixed this by giving every note participant a `Blogs` row,
|
||||||
|
> backfilling the IDs (none renumbered), making `ix_Blogs_BlogId` unique, and **dropping
|
||||||
|
> `BlogNames`**. There is no compatibility view: any query naming it fails with
|
||||||
|
> `no such table: BlogNames`. Two triggers now protect the IDs.
|
||||||
|
>
|
||||||
|
> **Porting an app:** replace `BlogNames` with `Blogs` everywhere. The columns you used,
|
||||||
|
> `BlogId` and `BlogName`, exist there with the same meaning. A name lookup
|
||||||
|
> (`SELECT BlogId FROM Blogs WHERE BlogName = ?`) is a primary-key probe, and an ID
|
||||||
|
> lookup or join (`JOIN Blogs b ON b.BlogId = n.NoteBlogId`) uses the unique
|
||||||
|
> `ix_Blogs_BlogId`. Every ID in `Notes` resolves to exactly one `Blogs` row. `Blogs.BlogId`
|
||||||
|
> is **no longer** a stale copy, so any code or docs that distrust it can drop that
|
||||||
|
> caveat. Never write `BlogId` or `BlogName` on a row that has an ID, and never delete such
|
||||||
|
> a row: the triggers reject all three. See [`Blogs`](#blogs).
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
## The three content tables
|
## The three content tables
|
||||||
|
|
||||||
| Table | Rows | What it is |
|
| Table | Rows | What it is |
|
||||||
|---|--:|---|
|
|---|--:|---|
|
||||||
| `Blogs` | 188,620 | The crawl registry — one row per known blog, plus crawl-state flags |
|
| `Blogs` | 198,560 | The crawl registry, one row per known blog plus crawl-state flags. Also the ID authority for blogs in `Notes` |
|
||||||
| `Posts` | 22,468 | Stored post content. Only 3,867 blogs actually have any |
|
| `Posts` | 22,468 | Stored post content. Only 3,867 blogs actually have any |
|
||||||
| `Notes` | 1,182,333 | The engagement graph: `NoteBlogId` acted on `(RootBlogId, PostID)` |
|
| `Notes` | 1,234,830 | The engagement graph: `NoteBlogId` acted on `(RootBlogId, PostID)` |
|
||||||
|
|
||||||
…supported by two lookup tables that exist only to keep `Notes` small:
|
(`Blogs` and `Notes` counts as of 2026-09-28; the rest as of 2026-08-07.)
|
||||||
|
|
||||||
|
…supported by one lookup table that exists only to keep `Notes` small:
|
||||||
|
|
||||||
| Table | Rows | What it is |
|
| Table | Rows | What it is |
|
||||||
|---|--:|---|
|
|---|--:|---|
|
||||||
| `BlogNames` | 20,430 | `BlogId` ⇄ `BlogName`. The ID authority for everything in `Notes` |
|
|
||||||
| `NoteTypes` | 5 | `TypeId` ⇄ `Type`. `like`, `reblog`, `reply`, `posted`, `post_attribution` |
|
| `NoteTypes` | 5 | `TypeId` ⇄ `Type`. `like`, `reblog`, `reply`, `posted`, `post_attribution` |
|
||||||
|
|
||||||
The engagement graph is the interesting part. 20,311 distinct blogs appear as engagers —
|
The engagement graph is the interesting part. 20,311 distinct blogs appear as engagers —
|
||||||
@@ -66,27 +88,47 @@ CREATE TABLE "Blogs" (
|
|||||||
PRIMARY KEY("BlogName")
|
PRIMARY KEY("BlogName")
|
||||||
);
|
);
|
||||||
|
|
||||||
CREATE INDEX ix_Blogs_BlogId ON Blogs (BlogId);
|
CREATE UNIQUE INDEX ix_Blogs_BlogId ON Blogs (BlogId);
|
||||||
|
|
||||||
|
CREATE TRIGGER trg_Blogs_BlogId_NoDelete -- no DELETE of a row that has a BlogId
|
||||||
|
CREATE TRIGGER trg_Blogs_BlogId_Immutable -- no change to its BlogId or BlogName
|
||||||
```
|
```
|
||||||
|
|
||||||
`BlogName` is the primary key, so it is the only indexed way in by name. There is no index
|
`BlogName` is the primary key, so it is the only indexed way in by name. There is no index
|
||||||
on any flag or date — filtering or sorting on those scans all 188k rows, which is
|
on any flag or date. Filtering or sorting on those scans the whole table, which is
|
||||||
affordable here and is not on `Notes`.
|
affordable here and is not on `Notes`.
|
||||||
|
|
||||||
**`BlogId` is new as of 2026-08-07 and is the join key to `Notes`.** It exists so that
|
**`BlogId` is the ID that `Notes.RootBlogId` and `Notes.NoteBlogId` store, and `Blogs` is
|
||||||
`Notes` can reach `Blogs` in a single integer hop rather than going through `BlogNames`
|
the only place it lives** (since 2026-09-28; see the banner at the top). The join to
|
||||||
and ending in a text comparison:
|
`Notes` is one integer hop on the unique index:
|
||||||
|
|
||||||
```sql
|
```sql
|
||||||
-- what you want
|
|
||||||
FROM Blogs B JOIN Notes N ON N.NoteBlogId = B.BlogId
|
FROM Blogs B JOIN Notes N ON N.NoteBlogId = B.BlogId
|
||||||
|
|
||||||
-- not this
|
|
||||||
FROM Blogs B JOIN BlogNames BN ON BN.BlogName = B.BlogName
|
|
||||||
JOIN Notes N ON N.NoteBlogId = BN.BlogId
|
|
||||||
```
|
```
|
||||||
|
|
||||||
**`BlogId` is NULL on 168,202 of 188,620 rows** — every blog that has never appeared in a
|
**Every blog that appears in `Notes` has a `Blogs` row with a `BlogId`.** `AddNote`
|
||||||
|
guarantees it through `RegisterBlog`, which runs in the note's own transaction:
|
||||||
|
|
||||||
|
```sql
|
||||||
|
INSERT OR IGNORE INTO Blogs (BlogName, DateAdded, DateModified, DateCreated)
|
||||||
|
VALUES (@name, @now, @now, @now);
|
||||||
|
UPDATE Blogs SET BlogId = (SELECT IFNULL(MAX(BlogId), 0) + 1 FROM Blogs)
|
||||||
|
WHERE BlogName = @name AND BlogId IS NULL;
|
||||||
|
```
|
||||||
|
|
||||||
|
- Unlike `AddBlog`, this does **not** skip names containing `deact`. A note by a
|
||||||
|
deactivated blog still needs an ID, so such blogs now get registry rows too, with the
|
||||||
|
usual defaults (`HasBeenOutput = 0`, `IsActive` left at its default).
|
||||||
|
- Assigning a `BlogId` is bookkeeping, so it **does not move `DateModified`**.
|
||||||
|
- `MAX(BlogId) + 1` is safe only because an ID can never be freed. The two triggers see
|
||||||
|
to that: deleting a row that has a `BlogId`, or changing its `BlogId` or `BlogName`,
|
||||||
|
aborts. Remove a blog with `IsActive = 0` instead. A blog renamed upstream gets a new
|
||||||
|
row. Rows with no `BlogId` can still be deleted or renamed freely.
|
||||||
|
- `INSERT OR REPLACE` on `Blogs` gets around the delete trigger (SQLite does not fire
|
||||||
|
delete triggers for REPLACE unless `recursive_triggers` is on), and it would wipe the
|
||||||
|
`BlogId`. It was already forbidden because it resets `IsActive`. Do not use it.
|
||||||
|
|
||||||
|
**`BlogId` is NULL on 165,887 of 198,560 rows**, every blog that has never appeared in a
|
||||||
note. That is the large majority, and it is not an error: the registry is far bigger than
|
note. That is the large majority, and it is not an error: the registry is far bigger than
|
||||||
the engagement graph. An inner join on `BlogId` therefore silently drops those blogs,
|
the engagement graph. An inner join on `BlogId` therefore silently drops those blogs,
|
||||||
which is usually what you want for engagement queries and is wrong for registry listings.
|
which is usually what you want for engagement queries and is wrong for registry listings.
|
||||||
@@ -154,6 +196,10 @@ Notable:
|
|||||||
on every row, and something has since started writing it. Anything that treated it as
|
on every row, and something has since started writing it. Anything that treated it as
|
||||||
permanently unset, or derived the type from post content instead, should be re-examined
|
permanently unset, or derived the type from post content instead, should be re-examined
|
||||||
against the live data. Rolodex still derives it.
|
against the live data. Rolodex still derives it.
|
||||||
|
- **`PostDate` is `yyyy-MM-dd HH:mm:ss GMT`** — UTC, as the Tumblr API sends it, unlike
|
||||||
|
the local-time `DateCreated`/`DateModified`. Text-file exports may carry RFC 1123
|
||||||
|
(`Fri, 14 Feb 2025 15:20:09 GMT`); every write path runs `PostDates.Normalize` to
|
||||||
|
convert it, and `../normalize-postdate.sql` fixed the 4 rows written before that.
|
||||||
- `HasImage = 1` on 14,026 rows. It records that the post *had* a picture, not that a
|
- `HasImage = 1` on 14,026 rows. It records that the post *had* a picture, not that a
|
||||||
usable URL was kept, so it is not a reliable predictor that anything will render.
|
usable URL was kept, so it is not a reliable predictor that anything will render.
|
||||||
- `PhotoURL` is largely unused; in practice the image markup lives inside `Body`.
|
- `PhotoURL` is largely unused; in practice the image markup lives inside `Body`.
|
||||||
@@ -183,8 +229,7 @@ CREATE INDEX ix_Notes_NoteBlogId ON Notes (NoteBlogId);
|
|||||||
|
|
||||||
**Integer IDs since 2026-08-07 — this is the breaking change.** `RootBlogName`,
|
**Integer IDs since 2026-08-07 — this is the breaking change.** `RootBlogName`,
|
||||||
`NoteBlogName` and `Type` are gone, replaced by `RootBlogId`, `NoteBlogId` and `TypeId`.
|
`NoteBlogName` and `Type` are gone, replaced by `RootBlogId`, `NoteBlogId` and `TypeId`.
|
||||||
Resolve them through [`BlogNames`](#blognames) and [`NoteTypes`](#notetypes), or join
|
Resolve blog IDs through `Blogs.BlogId` and types through [`NoteTypes`](#notetypes). The old names were text repeated on 1.18 million rows,
|
||||||
straight to `Blogs` on `BlogId`. The old names were text repeated on 1.18 million rows,
|
|
||||||
in the table *and* in every index over it; the swap took the file from 207 MB to 148 MB.
|
in the table *and* in every index over it; the swap took the file from 207 MB to 148 MB.
|
||||||
|
|
||||||
The **primary key column order is deliberately unchanged**, so the leading-prefix access
|
The **primary key column order is deliberately unchanged**, so the leading-prefix access
|
||||||
@@ -247,43 +292,21 @@ At 1.18M rows this is the table that dictates how the whole database has to be q
|
|||||||
- `replyText` is `'.'` on 1,167,464 rows — only `reply` notes carry real text. Those
|
- `replyText` is `'.'` on 1,167,464 rows — only `reply` notes carry real text. Those
|
||||||
dots are inherited from the old column default; new rows get `NULL` instead.
|
dots are inherited from the old column default; new rows get `NULL` instead.
|
||||||
|
|
||||||
**Resolve IDs by filtering the lookup, not by scanning `Notes`.** The lookup tables are
|
**Resolve IDs by filtering `Blogs`, not by scanning `Notes`.** A name predicate on `Blogs`
|
||||||
tiny and uniquely indexed, so pushing a name predicate into them costs nothing and lets
|
is a primary-key probe, so pushing it there costs nothing and lets the `Notes` index do
|
||||||
the `Notes` index do the work:
|
the work:
|
||||||
|
|
||||||
```sql
|
```sql
|
||||||
-- good: BlogNames resolves the name, then the index is searched
|
-- good: Blogs resolves the name, then the index is searched
|
||||||
SELECT * FROM Notes
|
SELECT * FROM Notes
|
||||||
WHERE NoteBlogId = (SELECT BlogId FROM BlogNames WHERE BlogName = ?);
|
WHERE NoteBlogId = (SELECT BlogId FROM Blogs WHERE BlogName = ?);
|
||||||
|
|
||||||
-- also good, same plan
|
-- also good, same plan
|
||||||
SELECT n.* FROM Notes n
|
SELECT n.* FROM Notes n
|
||||||
JOIN BlogNames b ON b.BlogId = n.NoteBlogId
|
JOIN Blogs b ON b.BlogId = n.NoteBlogId
|
||||||
WHERE b.BlogName = ?;
|
WHERE b.BlogName = ?;
|
||||||
```
|
```
|
||||||
|
|
||||||
### `BlogNames`
|
|
||||||
|
|
||||||
```sql
|
|
||||||
CREATE TABLE BlogNames (
|
|
||||||
BlogId INTEGER PRIMARY KEY,
|
|
||||||
BlogName TEXT NOT NULL UNIQUE
|
|
||||||
);
|
|
||||||
```
|
|
||||||
|
|
||||||
20,430 rows — every name appearing in `Notes` as either participant, and nothing else.
|
|
||||||
This is the **ID authority**: `Notes.RootBlogId` and `Notes.NoteBlogId` both point here,
|
|
||||||
and `Blogs.BlogId` is a copy of the value for the blogs that have one.
|
|
||||||
|
|
||||||
**12 of these names have no `Blogs` row.** The registry has never been a superset of the
|
|
||||||
engagement graph and still is not, so resolving an ID through `Blogs` rather than
|
|
||||||
`BlogNames` will occasionally find nothing. Use `BlogNames` when you need the name itself
|
|
||||||
and `Blogs` when you need registry columns.
|
|
||||||
|
|
||||||
IDs are assigned by SQLite and are **stable**: they are stored in over a million `Notes`
|
|
||||||
rows. Never renumber them. A blog that is renamed upstream should get a new row, not an
|
|
||||||
edit to an existing one, unless every `Notes` reference is migrated with it.
|
|
||||||
|
|
||||||
### `NoteTypes`
|
### `NoteTypes`
|
||||||
|
|
||||||
```sql
|
```sql
|
||||||
@@ -313,6 +336,9 @@ code to this table's contents, so prefer the join in anything long-lived.
|
|||||||
|
|
||||||
## Porting to the integer schema
|
## Porting to the integer schema
|
||||||
|
|
||||||
|
> Written for the 2026-08-07 change, and updated for 2026-09-28: wherever this section
|
||||||
|
> once said `BlogNames`, it now says `Blogs`. `BlogNames` no longer exists.
|
||||||
|
|
||||||
Everything here was checked against the live 148 MB file. There were 14 affected call
|
Everything here was checked against the live 148 MB file. There were 14 affected call
|
||||||
sites in `DataAccess.cs` and 16 in `RolodexRepository.cs`. TumblThree needs no changes —
|
sites in `DataAccess.cs` and 16 in `RolodexRepository.cs`. TumblThree needs no changes —
|
||||||
its single statement touches `Blogs.IsActive` and `BlogName` only.
|
its single statement touches `Blogs.IsActive` and `BlogName` only.
|
||||||
@@ -329,8 +355,8 @@ the result.
|
|||||||
|
|
||||||
| Was | Is now | Resolve via |
|
| Was | Is now | Resolve via |
|
||||||
|---|---|---|
|
|---|---|---|
|
||||||
| `Notes.RootBlogName` | `Notes.RootBlogId` | `BlogNames.BlogId` → `.BlogName` |
|
| `Notes.RootBlogName` | `Notes.RootBlogId` | `Blogs.BlogId` → `.BlogName` |
|
||||||
| `Notes.NoteBlogName` | `Notes.NoteBlogId` | `BlogNames.BlogId` → `.BlogName` |
|
| `Notes.NoteBlogName` | `Notes.NoteBlogId` | `Blogs.BlogId` → `.BlogName` |
|
||||||
| `Notes.Type` | `Notes.TypeId` | `NoteTypes.TypeId` → `.Type` |
|
| `Notes.Type` | `Notes.TypeId` | `NoteTypes.TypeId` → `.Type` |
|
||||||
| `ix_NoteBlogName01` | `ix_Notes_NoteBlogId` | — |
|
| `ix_NoteBlogName01` | `ix_Notes_NoteBlogId` | — |
|
||||||
|
|
||||||
@@ -343,14 +369,14 @@ the result.
|
|||||||
-- was
|
-- was
|
||||||
WHERE NoteBlogName = @Name
|
WHERE NoteBlogName = @Name
|
||||||
|
|
||||||
-- now, either form; both search ix_Notes_NoteBlogId after a unique-index lookup
|
-- now, either form; both search ix_Notes_NoteBlogId after a primary-key lookup
|
||||||
WHERE NoteBlogId = (SELECT BlogId FROM BlogNames WHERE BlogName = @Name)
|
WHERE NoteBlogId = (SELECT BlogId FROM Blogs WHERE BlogName = @Name)
|
||||||
-- or
|
-- or
|
||||||
JOIN BlogNames b ON b.BlogId = n.NoteBlogId WHERE b.BlogName = @Name
|
JOIN Blogs b ON b.BlogId = n.NoteBlogId WHERE b.BlogName = @Name
|
||||||
```
|
```
|
||||||
|
|
||||||
Measured 73 ms against 63 ms for the old text form on the busiest blog — the extra hop is
|
Measured at 73 ms against 63 ms for the old text form on the busiest blog, when the lookup
|
||||||
a unique-index probe on a 20k-row table and does not show.
|
was still `BlogNames`. The extra hop is one index probe and does not show.
|
||||||
|
|
||||||
### Joining `Notes` to `Blogs`
|
### Joining `Notes` to `Blogs`
|
||||||
|
|
||||||
@@ -360,12 +386,17 @@ This is the join to get right; it is the most common shape in both applications.
|
|||||||
-- was
|
-- was
|
||||||
FROM Blogs B INNER JOIN Notes N ON N.NoteBlogName = B.BlogName
|
FROM Blogs B INNER JOIN Notes N ON N.NoteBlogName = B.BlogName
|
||||||
|
|
||||||
-- now: one integer hop, using the new Blogs.BlogId
|
-- now: one integer hop, using Blogs.BlogId
|
||||||
FROM Blogs B INNER JOIN Notes N ON N.NoteBlogId = B.BlogId
|
FROM Blogs B INNER JOIN Notes N ON N.NoteBlogId = B.BlogId
|
||||||
```
|
```
|
||||||
|
|
||||||
Do **not** route this through `BlogNames` — that adds a hop and ends in the text
|
Joining `Notes` to `Posts` also goes through `Blogs`, since `Posts` has only a name:
|
||||||
comparison the change was meant to remove.
|
|
||||||
|
```sql
|
||||||
|
FROM Posts P
|
||||||
|
JOIN Blogs RB ON RB.BlogName = P.BlogName
|
||||||
|
JOIN Notes N ON N.RootBlogId = RB.BlogId AND N.PostID = P.PostID
|
||||||
|
```
|
||||||
|
|
||||||
### Selecting a name back out
|
### Selecting a name back out
|
||||||
|
|
||||||
@@ -374,12 +405,12 @@ comparison the change was meant to remove.
|
|||||||
SELECT NoteBlogName AS blogName, COUNT(*) FROM Notes ... GROUP BY NoteBlogName
|
SELECT NoteBlogName AS blogName, COUNT(*) FROM Notes ... GROUP BY NoteBlogName
|
||||||
|
|
||||||
-- now
|
-- now
|
||||||
SELECT bn.BlogName AS blogName, COUNT(*)
|
SELECT b.BlogName AS blogName, COUNT(*)
|
||||||
FROM Notes n JOIN BlogNames bn ON bn.BlogId = n.NoteBlogId
|
FROM Notes n JOIN Blogs b ON b.BlogId = n.NoteBlogId
|
||||||
... GROUP BY bn.BlogName
|
... GROUP BY b.BlogName
|
||||||
```
|
```
|
||||||
|
|
||||||
Group by `n.NoteBlogId` instead of `bn.BlogName` when you only need the name for display —
|
Group by `n.NoteBlogId` instead of `b.BlogName` when you only need the name for display —
|
||||||
grouping on the integer is cheaper and the name comes along for free.
|
grouping on the integer is cheaper and the name comes along for free.
|
||||||
|
|
||||||
### Filtering by type
|
### Filtering by type
|
||||||
@@ -403,27 +434,25 @@ because `TypeId` is `NOT NULL`.
|
|||||||
|
|
||||||
### Inserting a note
|
### Inserting a note
|
||||||
|
|
||||||
The crawler must ensure both names have IDs first. `INSERT OR IGNORE` on `BlogNames` is
|
The crawler must ensure both blogs have IDs first: run the `RegisterBlog` pair shown under
|
||||||
the whole of it — no read-back, no round trip, safe to run every time:
|
[`Blogs`](#blogs) for each name. No read-back, no round trip, and safe to run every time.
|
||||||
|
Then:
|
||||||
|
|
||||||
```sql
|
```sql
|
||||||
INSERT OR IGNORE INTO BlogNames (BlogName) VALUES (@rootBlogName);
|
|
||||||
INSERT OR IGNORE INTO BlogNames (BlogName) VALUES (@noteBlogName);
|
|
||||||
|
|
||||||
INSERT OR IGNORE INTO Notes
|
INSERT OR IGNORE INTO Notes
|
||||||
(RootBlogId, PostID, NoteBlogId, TimeStamp, TypeId,
|
(RootBlogId, PostID, NoteBlogId, TimeStamp, TypeId,
|
||||||
DatetimeCrawled, DateModified, DateCreated)
|
DatetimeCrawled, DateModified, DateCreated)
|
||||||
SELECT (SELECT BlogId FROM BlogNames WHERE BlogName = @rootBlogName),
|
SELECT (SELECT BlogId FROM Blogs WHERE BlogName = @rootBlogName),
|
||||||
@PostID,
|
@PostID,
|
||||||
(SELECT BlogId FROM BlogNames WHERE BlogName = @noteBlogName),
|
(SELECT BlogId FROM Blogs WHERE BlogName = @noteBlogName),
|
||||||
@TimeStamp,
|
@TimeStamp,
|
||||||
(SELECT TypeId FROM NoteTypes WHERE Type = @Type),
|
(SELECT TypeId FROM NoteTypes WHERE Type = @Type),
|
||||||
@DatetimeCrawled, @DateModified, @DateCreated;
|
@DatetimeCrawled, @DateModified, @DateCreated;
|
||||||
```
|
```
|
||||||
|
|
||||||
Verified: a genuinely new note inserts, and re-running the identical statement inserts 0.
|
Verified: a genuinely new note inserts, and re-running the identical statement inserts 0.
|
||||||
Run all three statements in one transaction so a crash cannot leave a name registered
|
Run the registrations and the insert in one transaction so a crash cannot leave a blog
|
||||||
with no note.
|
registered with no note.
|
||||||
|
|
||||||
**The duplicate-key error message has changed.** `DataAccess.cs` compares against the
|
**The duplicate-key error message has changed.** `DataAccess.cs` compares against the
|
||||||
literal string
|
literal string
|
||||||
@@ -448,7 +477,7 @@ on `TimeStamp` either before or after:
|
|||||||
```sql
|
```sql
|
||||||
-- now
|
-- now
|
||||||
UPDATE Notes SET replyText = @replyText, DateModified = @dateModified
|
UPDATE Notes SET replyText = @replyText, DateModified = @dateModified
|
||||||
WHERE NoteBlogId = (SELECT BlogId FROM BlogNames WHERE BlogName = @noteBlogName)
|
WHERE NoteBlogId = (SELECT BlogId FROM Blogs WHERE BlogName = @noteBlogName)
|
||||||
AND ABS(TimeStamp - @TimeStamp) <= 5
|
AND ABS(TimeStamp - @TimeStamp) <= 5
|
||||||
AND TypeId = (SELECT TypeId FROM NoteTypes WHERE Type = 'reply')
|
AND TypeId = (SELECT TypeId FROM NoteTypes WHERE Type = 'reply')
|
||||||
AND (replyText IS NULL OR replyText = '' OR replyText = '.')
|
AND (replyText IS NULL OR replyText = '' OR replyText = '.')
|
||||||
@@ -458,18 +487,15 @@ UPDATE Notes SET replyText = @replyText, DateModified = @dateModified
|
|||||||
Rolodex's soft-delete updates need no change beyond the `WHERE` clause — they set
|
Rolodex's soft-delete updates need no change beyond the `WHERE` clause — they set
|
||||||
`IsActive`, which is untouched.
|
`IsActive`, which is untouched.
|
||||||
|
|
||||||
### Three traps
|
### Two traps
|
||||||
|
|
||||||
**`Blogs.BlogId` is NULL on 168,202 of 188,620 rows.** Any inner join on it silently drops
|
**`Blogs.BlogId` is NULL on 165,887 of 198,560 rows.** Any inner join on it silently drops
|
||||||
every blog that has never appeared in a note. Correct for engagement queries; wrong for
|
every blog that has never appeared in a note. Correct for engagement queries; wrong for
|
||||||
registry listings, which need a `LEFT JOIN` or no join at all.
|
registry listings, which need a `LEFT JOIN` or no join at all.
|
||||||
|
|
||||||
**12 names in `BlogNames` have no `Blogs` row.** Resolving an ID to a name through `Blogs`
|
**IDs are stable and must stay so.** `Blogs.BlogId` and `NoteTypes.TypeId` are stored in
|
||||||
will occasionally find nothing. Use `BlogNames` for names and `Blogs` for registry columns.
|
over a million `Notes` rows. Never renumber. A blog renamed upstream gets a new row, not
|
||||||
|
an edited one. The `Blogs` triggers reject both.
|
||||||
**IDs are stable and must stay so.** `BlogNames.BlogId` and `NoteTypes.TypeId` are stored
|
|
||||||
in over a million `Notes` rows. Never renumber. A blog renamed upstream gets a new row,
|
|
||||||
not an edited one, unless every `Notes` reference migrates with it.
|
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
@@ -477,16 +503,12 @@ not an edited one, unless every `Notes` reference migrates with it.
|
|||||||
|
|
||||||
There are no foreign keys, and the tables do not perfectly agree:
|
There are no foreign keys, and the tables do not perfectly agree:
|
||||||
|
|
||||||
- 4 `Posts` rows name a blog with no `Blogs` row.
|
- 4 `Posts` rows name a blog with no `Blogs` row, so joins from `Posts` back to `Blogs`
|
||||||
- 12 of the 20,430 names in `BlogNames` have no `Blogs` row.
|
should tolerate a miss.
|
||||||
|
- `Notes` is covered: every `RootBlogId` and `NoteBlogId` resolves to a `Blogs` row.
|
||||||
So a name appearing in `Notes` or `Posts` is not a guarantee that the registry knows about
|
`retire-blognames.sql` checked this before committing, and `RegisterBlog` keeps it true.
|
||||||
it. Joins from those tables back to `Blogs` should tolerate a miss.
|
Before 2026-09-28, 12 to 17 note participants had no registry row. They now have stub
|
||||||
|
rows.
|
||||||
The integer schema does not fix this and was not meant to. `BlogNames` is deliberately
|
|
||||||
built from `Notes` rather than from `Blogs`, precisely so that the 12 unregistered
|
|
||||||
engagers keep their IDs and their rows. Had it been built from the registry, those notes
|
|
||||||
would have been dropped by the migration's inner joins.
|
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
@@ -538,11 +560,10 @@ Crawler bookkeeping. Rolodex ignores all of these.
|
|||||||
`DataAccess.cs` joins on it to decide what to collect:
|
`DataAccess.cs` joins on it to decide what to collect:
|
||||||
|
|
||||||
```sql
|
```sql
|
||||||
-- shape only; the ported GetBlogs joins Blogs directly on BlogId and needs no BlogNames hop
|
-- shape only
|
||||||
SELECT bn.BlogName, count(*)
|
SELECT b.BlogName, count(*)
|
||||||
FROM Notes n
|
FROM Notes n
|
||||||
JOIN Blogs b ON b.BlogId = n.NoteBlogId
|
JOIN Blogs b ON b.BlogId = n.NoteBlogId
|
||||||
JOIN BlogNames bn ON bn.BlogId = n.NoteBlogId
|
|
||||||
WHERE b.IsActive = @isActive AND ...
|
WHERE b.IsActive = @isActive AND ...
|
||||||
```
|
```
|
||||||
|
|
||||||
@@ -640,7 +661,7 @@ handled:
|
|||||||
SELECT 'Blogs', COUNT(*) FROM Blogs
|
SELECT 'Blogs', COUNT(*) FROM Blogs
|
||||||
UNION ALL SELECT 'Posts', COUNT(*) FROM Posts
|
UNION ALL SELECT 'Posts', COUNT(*) FROM Posts
|
||||||
UNION ALL SELECT 'Notes', COUNT(*) FROM Notes
|
UNION ALL SELECT 'Notes', COUNT(*) FROM Notes
|
||||||
UNION ALL SELECT 'BlogNames', COUNT(*) FROM BlogNames;
|
UNION ALL SELECT 'Blogs with a BlogId', COUNT(*) FROM Blogs WHERE BlogId IS NOT NULL;
|
||||||
|
|
||||||
-- note type mix (joins NoteTypes; Notes.Type no longer exists)
|
-- note type mix (joins NoteTypes; Notes.Type no longer exists)
|
||||||
SELECT t.Type, COUNT(*)
|
SELECT t.Type, COUNT(*)
|
||||||
@@ -664,8 +685,9 @@ SELECT COUNT(*) FROM (
|
|||||||
SELECT COUNT(*) FROM Posts p
|
SELECT COUNT(*) FROM Posts p
|
||||||
WHERE NOT EXISTS (SELECT 1 FROM Blogs b WHERE b.BlogName = p.BlogName);
|
WHERE NOT EXISTS (SELECT 1 FROM Blogs b WHERE b.BlogName = p.BlogName);
|
||||||
|
|
||||||
SELECT COUNT(*) FROM BlogNames bn
|
-- note participants with no Blogs.BlogId (expect 0; anything else is the pre-2026-09-28 drift)
|
||||||
WHERE NOT EXISTS (SELECT 1 FROM Blogs b WHERE b.BlogName = bn.BlogName);
|
SELECT COUNT(*) FROM (SELECT DISTINCT NoteBlogId AS Id FROM Notes) n
|
||||||
|
WHERE NOT EXISTS (SELECT 1 FROM Blogs b WHERE b.BlogId = n.Id);
|
||||||
|
|
||||||
-- space by object, to see where the file actually goes
|
-- space by object, to see where the file actually goes
|
||||||
SELECT name, SUM(pgsize)/1024/1024 AS mb
|
SELECT name, SUM(pgsize)/1024/1024 AS mb
|
||||||
|
|||||||
@@ -2,6 +2,10 @@
|
|||||||
-- Replaces the repeated blog-name and type TEXT in Notes with integer IDs.
|
-- Replaces the repeated blog-name and type TEXT in Notes with integer IDs.
|
||||||
-- Reduces TL.db from ~207 MB to ~148 MB (-29%).
|
-- Reduces TL.db from ~207 MB to ~148 MB (-29%).
|
||||||
--
|
--
|
||||||
|
-- SUPERSEDED IN PART, 2026-09-28: the BlogNames table this creates is no longer the
|
||||||
|
-- ID authority. Run retire-blognames.sql straight after this one; it moves the IDs
|
||||||
|
-- into Blogs.BlogId and drops BlogNames. The current app code assumes both have run.
|
||||||
|
--
|
||||||
-- THIS IS A BREAKING SCHEMA CHANGE. There is no compatibility layer. Every
|
-- THIS IS A BREAKING SCHEMA CHANGE. There is no compatibility layer. Every
|
||||||
-- query in URLNotesGrabberCORE and Rolodex that names Notes.RootBlogName,
|
-- query in URLNotesGrabberCORE and Rolodex that names Notes.RootBlogName,
|
||||||
-- Notes.NoteBlogName or Notes.Type stops working the moment this runs, and
|
-- Notes.NoteBlogName or Notes.Type stops working the moment this runs, and
|
||||||
|
|||||||
@@ -0,0 +1,31 @@
|
|||||||
|
-- normalize-postdate.sql
|
||||||
|
-- Rewrites Posts.PostDate values held in RFC 1123 form ("Fri, 14 Feb 2025 15:20:09 GMT")
|
||||||
|
-- into the column's canonical "yyyy-MM-dd HH:mm:ss GMT" (the Tumblr API's own format).
|
||||||
|
--
|
||||||
|
-- As of 2026-09-21 this matched 4 rows, all zombaee, from one text-file import. As text
|
||||||
|
-- they sort on the weekday name and never satisfy --fromDate / --toDate comparisons.
|
||||||
|
-- New writes are normalized in code by PostDates.Normalize, so this is a one-off.
|
||||||
|
--
|
||||||
|
-- Only PostDate changes. DateModified is left alone: the post content did not change.
|
||||||
|
--
|
||||||
|
-- HOW TO RUN: back up TL.db, then from the URLNotesGrabberCORE project folder:
|
||||||
|
-- sqlite3 TL.db < ../normalize-postdate.sql
|
||||||
|
|
||||||
|
SELECT 'before', COUNT(*) FROM Posts WHERE PostDate LIKE '___, __ ___ ____ __:__:__ GMT';
|
||||||
|
|
||||||
|
BEGIN;
|
||||||
|
UPDATE Posts
|
||||||
|
SET PostDate = substr(PostDate, 13, 4) || '-' ||
|
||||||
|
printf('%02d', (instr('JanFebMarAprMayJunJulAugSepOctNovDec', substr(PostDate, 9, 3)) + 2) / 3) || '-' ||
|
||||||
|
substr(PostDate, 6, 2) || ' ' ||
|
||||||
|
substr(PostDate, 18)
|
||||||
|
WHERE PostDate LIKE '___, __ ___ ____ __:__:__ GMT'
|
||||||
|
AND instr('JanFebMarAprMayJunJulAugSepOctNovDec', substr(PostDate, 9, 3)) % 3 = 1;
|
||||||
|
COMMIT;
|
||||||
|
|
||||||
|
-- VERIFY: expect 0, then a single shape '9999-99-99 99:99:99 GMT' (plus any NULL/blank)
|
||||||
|
SELECT 'after', COUNT(*) FROM Posts WHERE PostDate LIKE '___, %';
|
||||||
|
SELECT CASE WHEN PostDate GLOB '[0-9][0-9][0-9][0-9]-[0-9][0-9]-[0-9][0-9] [0-9][0-9]:[0-9][0-9]:[0-9][0-9] GMT'
|
||||||
|
THEN 'yyyy-MM-dd HH:mm:ss GMT' ELSE IFNULL(PostDate, '(null)') END AS shape,
|
||||||
|
COUNT(*)
|
||||||
|
FROM Posts GROUP BY 1 ORDER BY 2 DESC;
|
||||||
@@ -0,0 +1,143 @@
|
|||||||
|
-- retire-blognames.sql
|
||||||
|
-- Makes Blogs.BlogId the only ID authority for Notes and retires the BlogNames table.
|
||||||
|
--
|
||||||
|
-- WHY: normalize-notes.sql (2026-08-07) put the IDs in BlogNames and copied them into
|
||||||
|
-- Blogs.BlogId once. Nothing kept the copy current: by 2026-09-28, 12,238 blogs first seen
|
||||||
|
-- in a note after the migration had a BlogNames ID but Blogs.BlogId = NULL, so every query
|
||||||
|
-- joining Notes to Blogs on BlogId (GetBlogs and friends) silently skipped them -- 23,148
|
||||||
|
-- notes. Two copies of one ID drift; this leaves one.
|
||||||
|
--
|
||||||
|
-- What it does:
|
||||||
|
-- 1. Gives every BlogNames name a Blogs row (17 had none), carrying its ID over.
|
||||||
|
-- 2. Copies the ID onto every Blogs row that is missing it. IDs are never renumbered --
|
||||||
|
-- they are stored in 1.18M Notes rows.
|
||||||
|
-- 3. Proves every Notes ID resolves through Blogs before anything is dropped.
|
||||||
|
-- 4. Makes ix_Blogs_BlogId UNIQUE.
|
||||||
|
-- 5. Drops BlogNames. No compatibility view: any other app that still names it gets
|
||||||
|
-- "no such table: BlogNames" and must port to Blogs.BlogId (see TL.db.md).
|
||||||
|
-- 6. Adds triggers that stop a Blogs row holding a BlogId from being deleted, renamed or
|
||||||
|
-- renumbered -- the guarantees BlogNames gave by never being touched.
|
||||||
|
--
|
||||||
|
-- DateModified is NOT moved: assigning an ID is bookkeeping, not a content change. The 17
|
||||||
|
-- new stub rows get DateAdded/DateModified/DateCreated = now, as AddBlog would give them.
|
||||||
|
--
|
||||||
|
-- Runs after normalize-notes.sql. A backup from before 2026-08-07 needs both, in order.
|
||||||
|
--
|
||||||
|
-- HOW TO RUN:
|
||||||
|
-- 1. Stop every app that uses TL.db. Pause NextCloud sync.
|
||||||
|
-- 2. Back up TL.db: sqlite3 TL.db ".backup 'TL pre-retire-blognames.db'"
|
||||||
|
-- 3. sqlite3 -bail TL.db < retire-blognames.sql
|
||||||
|
-- -bail matters: a failed check aborts before COMMIT and nothing is changed.
|
||||||
|
-- In DB Browser, Execute SQL stops at the first error; then Revert Changes.
|
||||||
|
-- 4. Run the build of URLNotesGrabberCORE that no longer uses BlogNames. An older
|
||||||
|
-- build fails every AddNote with "no such table: BlogNames".
|
||||||
|
|
||||||
|
PRAGMA foreign_keys = off;
|
||||||
|
|
||||||
|
BEGIN;
|
||||||
|
|
||||||
|
-- Every check inserts one count here; the CHECK aborts the script on anything but 0.
|
||||||
|
CREATE TEMP TABLE MustBeZero (Check_ TEXT, n INTEGER CHECK (n = 0));
|
||||||
|
|
||||||
|
--------------------------------------------------------------------------
|
||||||
|
-- STEP 0: the two copies must not disagree anywhere they are both set
|
||||||
|
--------------------------------------------------------------------------
|
||||||
|
INSERT INTO MustBeZero
|
||||||
|
SELECT 'Blogs.BlogId differs from BlogNames', COUNT(*)
|
||||||
|
FROM Blogs b JOIN BlogNames bn ON bn.BlogName = b.BlogName
|
||||||
|
WHERE b.BlogId <> bn.BlogId;
|
||||||
|
|
||||||
|
INSERT INTO MustBeZero
|
||||||
|
SELECT 'Blogs.BlogId unknown to BlogNames', COUNT(*)
|
||||||
|
FROM Blogs b
|
||||||
|
WHERE b.BlogId IS NOT NULL
|
||||||
|
AND NOT EXISTS (SELECT 1 FROM BlogNames bn WHERE bn.BlogId = b.BlogId AND bn.BlogName = b.BlogName);
|
||||||
|
|
||||||
|
--------------------------------------------------------------------------
|
||||||
|
-- STEP 1: a Blogs row for every name Notes points at
|
||||||
|
--------------------------------------------------------------------------
|
||||||
|
-- No IsActive in the column list: it is not ours to write (defaults to live).
|
||||||
|
INSERT INTO Blogs (BlogName, DateAdded, DateModified, DateCreated, BlogId)
|
||||||
|
SELECT bn.BlogName,
|
||||||
|
strftime('%Y-%m-%d %H:%M:%S', 'now', 'localtime'),
|
||||||
|
strftime('%Y-%m-%d %H:%M:%S', 'now', 'localtime'),
|
||||||
|
strftime('%Y-%m-%d %H:%M:%S', 'now', 'localtime'),
|
||||||
|
bn.BlogId
|
||||||
|
FROM BlogNames bn
|
||||||
|
WHERE NOT EXISTS (SELECT 1 FROM Blogs b WHERE b.BlogName = bn.BlogName);
|
||||||
|
|
||||||
|
--------------------------------------------------------------------------
|
||||||
|
-- STEP 2: backfill the IDs Blogs never received
|
||||||
|
--------------------------------------------------------------------------
|
||||||
|
UPDATE Blogs
|
||||||
|
SET BlogId = (SELECT bn.BlogId FROM BlogNames bn WHERE bn.BlogName = Blogs.BlogName)
|
||||||
|
WHERE BlogId IS NULL
|
||||||
|
AND BlogName IN (SELECT BlogName FROM BlogNames);
|
||||||
|
|
||||||
|
--------------------------------------------------------------------------
|
||||||
|
-- STEP 3: prove Blogs now holds exactly what BlogNames held
|
||||||
|
--------------------------------------------------------------------------
|
||||||
|
INSERT INTO MustBeZero
|
||||||
|
SELECT 'BlogNames pair missing from Blogs', COUNT(*)
|
||||||
|
FROM BlogNames bn
|
||||||
|
WHERE NOT EXISTS (SELECT 1 FROM Blogs b WHERE b.BlogId = bn.BlogId AND b.BlogName = bn.BlogName);
|
||||||
|
|
||||||
|
INSERT INTO MustBeZero
|
||||||
|
SELECT 'Blogs IDs vs BlogNames rows', (SELECT COUNT(*) FROM Blogs WHERE BlogId IS NOT NULL) - (SELECT COUNT(*) FROM BlogNames);
|
||||||
|
|
||||||
|
INSERT INTO MustBeZero
|
||||||
|
SELECT 'Notes.RootBlogId unresolved', COUNT(*)
|
||||||
|
FROM (SELECT DISTINCT RootBlogId AS Id FROM Notes) n
|
||||||
|
WHERE NOT EXISTS (SELECT 1 FROM Blogs b WHERE b.BlogId = n.Id);
|
||||||
|
|
||||||
|
INSERT INTO MustBeZero
|
||||||
|
SELECT 'Notes.NoteBlogId unresolved', COUNT(*)
|
||||||
|
FROM (SELECT DISTINCT NoteBlogId AS Id FROM Notes) n
|
||||||
|
WHERE NOT EXISTS (SELECT 1 FROM Blogs b WHERE b.BlogId = n.Id);
|
||||||
|
|
||||||
|
--------------------------------------------------------------------------
|
||||||
|
-- STEP 4: one row per ID
|
||||||
|
--------------------------------------------------------------------------
|
||||||
|
-- UNIQUE still allows the NULLs on the ~168k blogs that have never appeared in a note.
|
||||||
|
DROP INDEX ix_Blogs_BlogId;
|
||||||
|
CREATE UNIQUE INDEX ix_Blogs_BlogId ON Blogs (BlogId);
|
||||||
|
|
||||||
|
--------------------------------------------------------------------------
|
||||||
|
-- STEP 5: BlogNames goes
|
||||||
|
--------------------------------------------------------------------------
|
||||||
|
DROP TABLE BlogNames;
|
||||||
|
|
||||||
|
--------------------------------------------------------------------------
|
||||||
|
-- STEP 6: what BlogNames guaranteed by never being written
|
||||||
|
--------------------------------------------------------------------------
|
||||||
|
-- A deleted row would orphan its notes, and MAX(BlogId) + 1 in RegisterBlog could then
|
||||||
|
-- hand the same ID to a different blog. Remove a blog with IsActive = 0 instead.
|
||||||
|
CREATE TRIGGER trg_Blogs_BlogId_NoDelete
|
||||||
|
BEFORE DELETE ON Blogs
|
||||||
|
WHEN OLD.BlogId IS NOT NULL
|
||||||
|
BEGIN
|
||||||
|
SELECT RAISE(ABORT, 'Blogs row has a BlogId that Notes points at; set IsActive = 0 instead of deleting');
|
||||||
|
END;
|
||||||
|
|
||||||
|
-- A blog renamed upstream is a new blog to Tumblr's API and gets a new row. Editing the name
|
||||||
|
-- in place would re-attribute every note to it; changing the ID would orphan them.
|
||||||
|
CREATE TRIGGER trg_Blogs_BlogId_Immutable
|
||||||
|
BEFORE UPDATE OF BlogId, BlogName ON Blogs
|
||||||
|
WHEN OLD.BlogId IS NOT NULL
|
||||||
|
AND (NEW.BlogId IS NOT OLD.BlogId OR NEW.BlogName IS NOT OLD.BlogName)
|
||||||
|
BEGIN
|
||||||
|
SELECT RAISE(ABORT, 'BlogId and BlogName are fixed once a blog has a BlogId; Notes rows point at it');
|
||||||
|
END;
|
||||||
|
|
||||||
|
DROP TABLE temp.MustBeZero;
|
||||||
|
|
||||||
|
COMMIT;
|
||||||
|
|
||||||
|
--------------------------------------------------------------------------
|
||||||
|
-- VERIFY
|
||||||
|
--------------------------------------------------------------------------
|
||||||
|
-- SELECT COUNT(*) FROM sqlite_master WHERE name = 'BlogNames'; -- expect: 0
|
||||||
|
-- SELECT COUNT(*) FROM Blogs WHERE BlogId IS NOT NULL; -- expect: the old BlogNames row count
|
||||||
|
-- SELECT sql FROM sqlite_master WHERE name = 'ix_Blogs_BlogId'; -- expect: CREATE UNIQUE INDEX
|
||||||
|
-- SELECT name FROM sqlite_master WHERE type = 'trigger'; -- expect: both triggers
|
||||||
|
-- PRAGMA integrity_check; -- expect: ok
|
||||||
+32
-12
@@ -82,7 +82,8 @@ WITH expected(tbl, col, alter_stmt) AS (
|
|||||||
-- Blogs.BlogId (2026-08-07) is the single-hop join key into Notes. Deliberately NOT
|
-- Blogs.BlogId (2026-08-07) is the single-hop join key into Notes. Deliberately NOT
|
||||||
-- auto-fixable: an added-but-empty BlogId makes every engagement join return zero
|
-- auto-fixable: an added-but-empty BlogId makes every engagement join return zero
|
||||||
-- rows silently, which is worse than the hard error a missing column gives.
|
-- rows silently, which is worse than the hard error a missing column gives.
|
||||||
('Blogs','BlogId', 'MANUAL REVIEW - see query 1d: run normalize-notes.sql'),
|
-- Since 2026-09-28 it is the only blog-ID authority (query 1e).
|
||||||
|
('Blogs','BlogId', 'MANUAL REVIEW - see queries 1d/1e: run normalize-notes.sql, then retire-blognames.sql'),
|
||||||
|
|
||||||
-- Notes (base columns: manual review if missing)
|
-- Notes (base columns: manual review if missing)
|
||||||
-- Integer IDs since 2026-08-07. RootBlogName/NoteBlogName/Type are GONE, not renamed
|
-- Integer IDs since 2026-08-07. RootBlogName/NoteBlogName/Type are GONE, not renamed
|
||||||
@@ -101,11 +102,11 @@ WITH expected(tbl, col, alter_stmt) AS (
|
|||||||
-- placeholder. EnsureReplyTextColumnExists in DataAccess.cs adds it the same way.
|
-- placeholder. EnsureReplyTextColumnExists in DataAccess.cs adds it the same way.
|
||||||
('Notes','replyText', 'ALTER TABLE Notes ADD COLUMN replyText TEXT;'),
|
('Notes','replyText', 'ALTER TABLE Notes ADD COLUMN replyText TEXT;'),
|
||||||
|
|
||||||
-- BlogNames / NoteTypes (the lookup tables Notes resolves its IDs through, 2026-08-07).
|
-- NoteTypes (the lookup table Notes resolves TypeId through, 2026-08-07).
|
||||||
-- Not auto-fixable: an empty BlogNames does not mean "add the table", it means the
|
-- Not auto-fixable: an empty NoteTypes does not mean "add the table", it means the
|
||||||
-- Notes rows have nothing to resolve against. Rebuild with normalize-notes.sql.
|
-- Notes rows have nothing to resolve against. Rebuild with normalize-notes.sql.
|
||||||
('BlogNames','BlogId', 'MANUAL REVIEW - see query 1d: run normalize-notes.sql'),
|
-- BlogNames is not listed: it was dropped on 2026-09-28. Query 1e reports a file
|
||||||
('BlogNames','BlogName', 'MANUAL REVIEW - see query 1d: run normalize-notes.sql'),
|
-- that still has it.
|
||||||
('NoteTypes','TypeId', 'MANUAL REVIEW - see query 1d: run normalize-notes.sql'),
|
('NoteTypes','TypeId', 'MANUAL REVIEW - see query 1d: run normalize-notes.sql'),
|
||||||
('NoteTypes','Type', 'MANUAL REVIEW - see query 1d: run normalize-notes.sql'),
|
('NoteTypes','Type', 'MANUAL REVIEW - see query 1d: run normalize-notes.sql'),
|
||||||
|
|
||||||
@@ -123,7 +124,6 @@ actual(tbl, col) AS (
|
|||||||
SELECT 'Posts', name FROM pragma_table_info('Posts')
|
SELECT 'Posts', name FROM pragma_table_info('Posts')
|
||||||
UNION ALL SELECT 'Blogs', name FROM pragma_table_info('Blogs')
|
UNION ALL SELECT 'Blogs', name FROM pragma_table_info('Blogs')
|
||||||
UNION ALL SELECT 'Notes', name FROM pragma_table_info('Notes')
|
UNION ALL SELECT 'Notes', name FROM pragma_table_info('Notes')
|
||||||
UNION ALL SELECT 'BlogNames', name FROM pragma_table_info('BlogNames')
|
|
||||||
UNION ALL SELECT 'NoteTypes', name FROM pragma_table_info('NoteTypes')
|
UNION ALL SELECT 'NoteTypes', name FROM pragma_table_info('NoteTypes')
|
||||||
UNION ALL SELECT 'DailyAPICount', name FROM pragma_table_info('DailyAPICount')
|
UNION ALL SELECT 'DailyAPICount', name FROM pragma_table_info('DailyAPICount')
|
||||||
UNION ALL SELECT 'ApiKeyPoolState', name FROM pragma_table_info('ApiKeyPoolState')
|
UNION ALL SELECT 'ApiKeyPoolState', name FROM pragma_table_info('ApiKeyPoolState')
|
||||||
@@ -146,7 +146,7 @@ ORDER BY (e.alter_stmt LIKE 'ALTER%') DESC, e.tbl, e.col;
|
|||||||
-- 1b. MISSING TABLES: expected tables that don't exist at all in this DB.
|
-- 1b. MISSING TABLES: expected tables that don't exist at all in this DB.
|
||||||
-- Zero rows = good.
|
-- Zero rows = good.
|
||||||
WITH expected_tables(tbl) AS (
|
WITH expected_tables(tbl) AS (
|
||||||
VALUES ('Posts'),('Blogs'),('Notes'),('BlogNames'),('NoteTypes'),('DailyAPICount'),
|
VALUES ('Posts'),('Blogs'),('Notes'),('NoteTypes'),('DailyAPICount'),
|
||||||
('ApiKeyPoolState'),('ApiKeyPoolMeta')
|
('ApiKeyPoolState'),('ApiKeyPoolMeta')
|
||||||
)
|
)
|
||||||
SELECT et.tbl AS missing_table
|
SELECT et.tbl AS missing_table
|
||||||
@@ -180,7 +180,6 @@ WITH expected(tbl, col) AS (
|
|||||||
('Notes','RootBlogId'),('Notes','PostID'),('Notes','NoteBlogId'),('Notes','TimeStamp'),
|
('Notes','RootBlogId'),('Notes','PostID'),('Notes','NoteBlogId'),('Notes','TimeStamp'),
|
||||||
('Notes','TypeId'),('Notes','DatetimeCrawled'),('Notes','DateModified'),('Notes','DateCreated'),
|
('Notes','TypeId'),('Notes','DatetimeCrawled'),('Notes','DateModified'),('Notes','DateCreated'),
|
||||||
('Notes','replyText'),('Notes','IsActive'),
|
('Notes','replyText'),('Notes','IsActive'),
|
||||||
('BlogNames','BlogId'),('BlogNames','BlogName'),
|
|
||||||
('NoteTypes','TypeId'),('NoteTypes','Type'),
|
('NoteTypes','TypeId'),('NoteTypes','Type'),
|
||||||
('DailyAPICount','Date'),('DailyAPICount','APICount'),
|
('DailyAPICount','Date'),('DailyAPICount','APICount'),
|
||||||
('ApiKeyPoolState','KeyName'),('ApiKeyPoolState','RetryUntil'),
|
('ApiKeyPoolState','KeyName'),('ApiKeyPoolState','RetryUntil'),
|
||||||
@@ -190,7 +189,6 @@ actual(tbl, col) AS (
|
|||||||
SELECT 'Posts', name FROM pragma_table_info('Posts')
|
SELECT 'Posts', name FROM pragma_table_info('Posts')
|
||||||
UNION ALL SELECT 'Blogs', name FROM pragma_table_info('Blogs')
|
UNION ALL SELECT 'Blogs', name FROM pragma_table_info('Blogs')
|
||||||
UNION ALL SELECT 'Notes', name FROM pragma_table_info('Notes')
|
UNION ALL SELECT 'Notes', name FROM pragma_table_info('Notes')
|
||||||
UNION ALL SELECT 'BlogNames', name FROM pragma_table_info('BlogNames')
|
|
||||||
UNION ALL SELECT 'NoteTypes', name FROM pragma_table_info('NoteTypes')
|
UNION ALL SELECT 'NoteTypes', name FROM pragma_table_info('NoteTypes')
|
||||||
UNION ALL SELECT 'DailyAPICount', name FROM pragma_table_info('DailyAPICount')
|
UNION ALL SELECT 'DailyAPICount', name FROM pragma_table_info('DailyAPICount')
|
||||||
UNION ALL SELECT 'ApiKeyPoolState', name FROM pragma_table_info('ApiKeyPoolState')
|
UNION ALL SELECT 'ApiKeyPoolState', name FROM pragma_table_info('ApiKeyPoolState')
|
||||||
@@ -209,7 +207,7 @@ ORDER BY a.tbl, a.col;
|
|||||||
--
|
--
|
||||||
-- This is the one failure SECTION 2 cannot fix. Notes.RootBlogName /
|
-- This is the one failure SECTION 2 cannot fix. Notes.RootBlogName /
|
||||||
-- NoteBlogName / Type were replaced by RootBlogId / NoteBlogId / TypeId
|
-- NoteBlogName / Type were replaced by RootBlogId / NoteBlogId / TypeId
|
||||||
-- resolving through BlogNames and NoteTypes -- a data migration, not an
|
-- resolving through (then) BlogNames and NoteTypes -- a data migration, not an
|
||||||
-- ADD COLUMN. There is no compatibility view, so the current code fails
|
-- ADD COLUMN. There is no compatibility view, so the current code fails
|
||||||
-- outright ("no such column: RootBlogId") against such a file.
|
-- outright ("no such column: RootBlogId") against such a file.
|
||||||
--
|
--
|
||||||
@@ -223,6 +221,28 @@ WHERE lower(name) IN ('rootblogname','noteblogname','type')
|
|||||||
HAVING COUNT(*) > 0;
|
HAVING COUNT(*) > 0;
|
||||||
|
|
||||||
|
|
||||||
|
-- 1e. BLOGNAMES NOT RETIRED: a backup from between 2026-08-07 and 2026-09-28, when
|
||||||
|
-- BlogNames still held the IDs and Blogs.BlogId was an
|
||||||
|
-- unmaintained copy. Zero rows = good.
|
||||||
|
--
|
||||||
|
-- The current code resolves every Notes ID through Blogs.BlogId and never writes
|
||||||
|
-- BlogNames, so against such a file new blogs get IDs that can collide with
|
||||||
|
-- BlogNames' and every blog missing from Blogs.BlogId stays invisible to GetBlogs.
|
||||||
|
--
|
||||||
|
-- Fix: back up, then run retire-blognames.sql (after normalize-notes.sql if 1d
|
||||||
|
-- also reported). It checks itself and changes nothing if a check fails.
|
||||||
|
SELECT 'BlogNames still exists (' || type || ') -- run retire-blognames.sql' AS blognames_not_retired
|
||||||
|
FROM sqlite_master
|
||||||
|
WHERE lower(name) = 'blognames'
|
||||||
|
UNION ALL
|
||||||
|
SELECT 'Blogs.BlogId is not UNIQUE -- run retire-blognames.sql'
|
||||||
|
WHERE NOT EXISTS (SELECT 1 FROM pragma_index_list('Blogs') WHERE name = 'ix_Blogs_BlogId' AND "unique" = 1)
|
||||||
|
UNION ALL
|
||||||
|
SELECT 'BlogId guard trigger missing: ' || t.name || ' -- run retire-blognames.sql'
|
||||||
|
FROM (SELECT 'trg_Blogs_BlogId_NoDelete' AS name UNION ALL SELECT 'trg_Blogs_BlogId_Immutable') t
|
||||||
|
WHERE NOT EXISTS (SELECT 1 FROM sqlite_master m WHERE m.type = 'trigger' AND m.name = t.name);
|
||||||
|
|
||||||
|
|
||||||
-- ============================================================================
|
-- ============================================================================
|
||||||
-- SECTION 2 -- FIX (opt-in, additive only)
|
-- SECTION 2 -- FIX (opt-in, additive only)
|
||||||
--
|
--
|
||||||
@@ -233,8 +253,8 @@ HAVING COUNT(*) > 0;
|
|||||||
-- subset. These are the 8 additive migration columns and nothing else; the
|
-- subset. These are the 8 additive migration columns and nothing else; the
|
||||||
-- likes high-water-mark reset is intentionally NOT included.
|
-- likes high-water-mark reset is intentionally NOT included.
|
||||||
--
|
--
|
||||||
-- Nothing here addresses query 1d. The Notes integer schema is a data migration
|
-- Nothing here addresses queries 1d or 1e. Those are data migrations
|
||||||
-- (normalize-notes.sql) and cannot be reached by adding columns.
|
-- (normalize-notes.sql, retire-blognames.sql) and cannot be reached by adding columns.
|
||||||
-- ============================================================================
|
-- ============================================================================
|
||||||
|
|
||||||
-- ALTER TABLE Posts ADD COLUMN PostType TEXT;
|
-- ALTER TABLE Posts ADD COLUMN PostType TEXT;
|
||||||
|
|||||||
Reference in New Issue
Block a user