Compare commits
10
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
a28c5cc9ec | ||
|
|
d6a96f7885 | ||
|
|
864b468d96 | ||
|
|
e6a3efba5b | ||
|
|
34da632e6a | ||
|
|
954ec353a5 | ||
|
|
c9530a3718 | ||
|
|
3d61cbb6ea | ||
|
|
0d9641db08 | ||
|
|
b31d5842cc |
@@ -31,7 +31,7 @@ dotnet run -- --test [blogname] [postID] # Test API for specific post
|
||||
- `--test [blogname] [postID]`: Test API note collection
|
||||
- `--posts`: Export post blogs to file
|
||||
- `--blogs`: Export blog list to file
|
||||
- `--collect`: Collect notes for all posts in DB
|
||||
- `--collect [0|1] [datetime] [blogname]`: Collect notes for posts in DB. Optional `blogname` restricts the run to one blog (exact match), e.g. `--collect 1 zomb-eh`
|
||||
- `--blogsR`: Export reply blogs to file
|
||||
- `--blogsO [start] [stop]`: Export blogs within range
|
||||
|
||||
|
||||
@@ -47,6 +47,36 @@ say nothing about the item being fetched, so they must not be recorded as per-it
|
||||
- Long-running commands return exit 3 when a pass ends incomplete (rate-limit pause, breaker trip, or
|
||||
skipped items), so a caller can distinguish that from a clean run
|
||||
|
||||
### `Notes` Stores Integer IDs, Not Names
|
||||
As of 2026-08-07 `Notes.RootBlogName`, `NoteBlogName` and `Type` are gone, replaced by
|
||||
`RootBlogId`, `NoteBlogId` and `TypeId` resolving through the `BlogNames` and `NoteTypes`
|
||||
lookup tables. There is no compatibility view — naming an old column is a hard SQLite
|
||||
error, so unlike `IsActive` this is a hard cut with no runtime probe. Full detail in
|
||||
`URLNotesGrabberCORE/TL.db.md`.
|
||||
|
||||
- **Joining `Notes` to `Blogs` goes through `Blogs.BlogId`**, not `BlogNames`:
|
||||
`FROM Blogs B INNER JOIN Notes N ON N.NoteBlogId = B.BlogId`. Routing it through
|
||||
`BlogNames` adds a hop and ends in the text comparison the migration removed
|
||||
- **Joining `Notes` to `Posts` is the opposite** — `Posts` has only `BlogName`, so it must
|
||||
go through `BlogNames` (`GetRepliesWithFilledText`). This is the only such join
|
||||
- **Resolve a name by filtering the lookup, never by scanning `Notes`**:
|
||||
`WHERE NoteBlogId = (SELECT BlogId FROM BlogNames WHERE BlogName = @name)`. The subquery
|
||||
is a unique-index probe on 20k rows and does not show against the 1.18M-row table
|
||||
- **`AddNote` registers both blog names *and* the note type** with `INSERT OR IGNORE`
|
||||
before inserting, all in one transaction. `NoteTypes` is a table rather than a `CHECK`
|
||||
constraint precisely so an unseen type is an `INSERT`; without that registration it
|
||||
would resolve to `NULL` and fail the `NOT NULL` on `TypeId`, losing the note
|
||||
- **`Blogs.BlogId` is NULL on 168,202 of 188,620 rows** — every blog that has never
|
||||
appeared in a note. An inner join on it silently drops them. Correct for engagement
|
||||
queries, wrong for anything listing the registry
|
||||
- **IDs are stable and must never be renumbered.** They are stored in 1.18M `Notes` rows.
|
||||
A blog renamed upstream gets a new `BlogNames` row, not an edited one
|
||||
- Prefer `TypeId = (SELECT TypeId FROM NoteTypes WHERE Type = 'reply')` over a hardcoded
|
||||
ID. A negated `TypeId NOT IN (SELECT …)` is only correct because `TypeId` is `NOT NULL`
|
||||
- Duplicate-key detection uses `IsNotesDuplicateKey`, which matches the constraint and the
|
||||
table rather than an exact column list. The old literal string comparison broke silently
|
||||
on this rename — do not reintroduce one
|
||||
|
||||
### `IsActive` Is Not Ours To Write
|
||||
`Blogs.IsActive`, `Posts.IsActive` and `Notes.IsActive` are removal flags set by other tools
|
||||
(Rolodex). `0` means removed; anything else, including `NULL`, means live. Full detail in
|
||||
|
||||
+109
-100
@@ -1,100 +1,109 @@
|
||||
<?xml version="1.0" encoding="UTF-8"?><sqlb_project><db path="D:/NextCloud/C#/URLNotesGrabberCORE/URLNotesGrabberCORE/TL.db" readonly="0" foreign_keys="1" case_sensitive_like="0" temp_store="0" wal_autocheckpoint="1000" synchronous="2"/><attached/><window><main_tabs open="structure browser pragmas query" current="3"/></window><tab_structure><column_width id="0" width="300"/><column_width id="1" width="0"/><column_width id="2" width="100"/><column_width id="3" width="4486"/><column_width id="4" width="0"/><expanded_item id="0" parent="1"/><expanded_item id="1" parent="1"/><expanded_item id="2" parent="1"/><expanded_item id="3" parent="1"/></tab_structure><tab_browse><table title="Posts" custom_title="0" dock_id="4" table="4,5:mainPosts"/><dock_state state="000000ff00000000fd0000000100000002000005470000029afc0100000006fb000000160064006f0063006b00420072006f00770073006500310100000000000004a10000000000000000fb000000160064006f0063006b00420072006f00770073006500320100000000000004a10000000000000000fb000000160064006f0063006b00420072006f00770073006500330100000000000004a10000000000000000fb000000160064006f0063006b00420072006f00770073006500350100000000000005f40000000000000000fb000000160064006f0063006b00420072006f00770073006500340100000000000005470000011100fffffffb000000160064006f0063006b00420072006f00770073006500340100000000000005f40000000000000000000005470000000000000004000000040000000800000008fc00000000"/><default_encoding codec=""/><browse_table_settings><table schema="main" name="ApiKeyPoolMeta" show_row_id="0" encoding="" plot_x_axis="" unlock_view_pk="_rowid_" freeze_columns="0"><sort/><column_widths><column index="1" value="29"/><column index="2" value="64"/></column_widths><filter_values/><conditional_formats/><row_id_formats/><display_formats/><hidden_columns/><plot_y_axes/><global_filter/></table><table schema="main" name="Blogs" show_row_id="0" encoding="" plot_x_axis="" unlock_view_pk="_rowid_" freeze_columns="0"><sort><column index="7" mode="1"/></sort><column_widths><column index="1" value="257"/><column index="2" value="108"/><column index="3" value="63"/><column index="4" value="156"/><column index="5" value="60"/><column index="6" value="81"/><column index="7" value="85"/><column index="8" value="156"/><column index="9" value="156"/><column index="10" value="151"/><column index="11" value="125"/><column index="12" value="129"/></column_widths><filter_values><column index="4" value="1"/><column index="7" value=">2026-05-27 17:22:36"/></filter_values><conditional_formats/><row_id_formats/><display_formats/><hidden_columns/><plot_y_axes/><global_filter/></table><table schema="main" name="Notes" show_row_id="0" encoding="" plot_x_axis="" unlock_view_pk="_rowid_" freeze_columns="0"><sort/><column_widths><column index="1" value="198"/><column index="2" value="144"/><column index="3" value="251"/><column index="4" value="84"/><column index="5" value="53"/><column index="6" value="300"/><column index="7" value="116"/><column index="8" value="152"/><column index="9" value="89"/><column index="10" value="63"/></column_widths><filter_values/><conditional_formats/><row_id_formats/><display_formats/><hidden_columns/><plot_y_axes/><global_filter/></table><table schema="main" name="Posts" show_row_id="0" encoding="" plot_x_axis="" unlock_view_pk="_rowid_" freeze_columns="0"><sort><column index="14" mode="1"/></sort><column_widths><column index="1" value="236"/><column index="2" value="144"/><column index="3" value="126"/><column index="4" value="32"/><column index="5" value="32"/><column index="6" value="32"/><column index="7" value="32"/><column index="8" value="32"/><column index="9" value="0"/><column index="10" value="0"/><column index="11" value="0"/><column index="12" value="243"/><column index="13" value="300"/><column index="14" value="53"/><column index="15" value="37351"/><column index="16" value="300"/><column index="17" value="41"/><column index="18" value="75"/><column index="19" value="96"/><column index="20" value="300"/><column index="21" value="96"/><column index="22" value="300"/><column index="23" value="300"/><column index="24" value="548"/><column index="25" value="60"/><column index="26" value="213"/><column index="27" value="532"/><column index="28" value="152"/><column index="29" value="152"/><column index="30" value="69"/><column index="31" value="63"/></column_widths><filter_values><column index="2" value="0"/><column index="1" value="734568371821084672"/></filter_values><conditional_formats/><row_id_formats/><display_formats/><hidden_columns><column index="9" value="1"/><column index="10" value="1"/><column index="11" value="1"/></hidden_columns><plot_y_axes/><global_filter/></table></browse_table_settings></tab_browse><tab_sql><sql name="SQL 1">UPDATE Posts
|
||||
SET HasNotesGathered = 0
|
||||
WHERE (BlogName, PostID) IN (
|
||||
SELECT p.BlogName, p.PostID
|
||||
FROM Posts p
|
||||
WHERE p.HasNotesGathered = 1
|
||||
AND P.notesGatheredDatetime < 1774294520
|
||||
AND EXISTS (
|
||||
SELECT 1
|
||||
FROM Notes n
|
||||
WHERE n.PostID = p.PostID
|
||||
AND n.RootBlogName = p.BlogName
|
||||
--AND n.Type NOT IN ('reblog', 'reply')
|
||||
)
|
||||
ORDER BY P.PostDate ASC
|
||||
--LIMIT 500
|
||||
);</sql><sql name="Mark Blogs">select *
|
||||
from Blogs
|
||||
--update blogs set HasBeenOutput = 1
|
||||
where HasBeenOutput = 0
|
||||
AND
|
||||
blogname in
|
||||
(
|
||||
'teaberrybee',
|
||||
'reddevilgoddesstoo',
|
||||
'waywardog13',
|
||||
'wzjustbrowsing-blog',
|
||||
'lewerta',
|
||||
'nudenymph',
|
||||
'caylachief'
|
||||
|
||||
)</sql><sql name="New Notes">select P.slug, N.replyText, n.RootBlogName, n.PostID, NoteBlogName || '.tumblr.com' as NoteBlogName, DatetimeCrawled, TimeStamp, type, n.RootBlogName || '.tumblr.com/post/' || n.postid, datetime(timestamp, 'unixepoch')
|
||||
from Notes N inner join Posts P on p.PostID = n.PostID
|
||||
where
|
||||
DatetimeCrawled > '2026-08-07 11:47:22' and type like 'r%'
|
||||
and P.IsActive = 1
|
||||
order by n.DatetimeCrawled</sql><sql name="Pull Blogs">SELECT distinct
|
||||
'''' || blogname || ''',',
|
||||
blogs.*
|
||||
, blogname || '.tumblr.com'
|
||||
FROM
|
||||
Blogs
|
||||
inner JOIN
|
||||
Notes on notes.noteBlogName = blogs.BlogName
|
||||
WHERE
|
||||
HasBeenOutput = 0 and type = 'reblog'
|
||||
order by
|
||||
Notes.Type desc,
|
||||
DateAdded desc
|
||||
LIMIT 100;</sql><sql name="SQL 7">WITH ReplyCounts AS (
|
||||
SELECT
|
||||
NoteBlogName,
|
||||
COUNT(DISTINCT replyText) AS DistinctReplyCount
|
||||
FROM Notes
|
||||
where replyText <> '.'
|
||||
GROUP BY NoteBlogName
|
||||
)
|
||||
SELECT
|
||||
n.RootBlogName || '.tumblr.com/post/' || n.PostID AS PostURL, postid,
|
||||
n.NoteBlogName,
|
||||
n.replyText,
|
||||
c.DistinctReplyCount
|
||||
FROM Notes n
|
||||
JOIN ReplyCounts c ON n.NoteBlogName = c.NoteBlogName
|
||||
where replyText <> '.' and type <> 'reply'
|
||||
--AND N.NoteBlogName NOT IN ( 'roadblocker21', 'thesaddemon666', 'edwardabbeyhoffman', 'tattedsoldier20', 'zomb-eh', 'animalistic13', 'indken', 'maccloud1592',
|
||||
--'moss-wizard', 'supertrucker12682', 'exploringthrupics', 'padeyepete' )
|
||||
order by c.DistinctReplyCount desc, n.NoteBlogName, n.DateModified desc, replyText, RootBlogName, PostID</sql><sql name="Collect">WITH PostsWithCount AS ( SELECT P.BlogName, P.PostID, 1925013599 AS LatestNoteTimestamp, P.NotesGatheredDateTime, COUNT(P.PostID) OVER(PARTITION BY P.BlogName) AS CNT, P.HasNotesGathered, P.NotFound, P.PostDate FROM Posts P WHERE COALESCE(P.IsActive, 1) = 1 ), Unioned AS ( SELECT BlogName, PostID, LatestNoteTimestamp, NotesGatheredDateTime, CNT, PostDate FROM PostsWithCount WHERE NotFound = 0 AND HasNotesGathered = 0 UNION SELECT BlogName, PostID, LatestNoteTimestamp, NotesGatheredDateTime, CNT, PostDate FROM PostsWithCount WHERE BlogName = 'zomb-eh' AND NotFound = 0 AND NotesGatheredDateTime < unixepoch('now', 'localtime', '-3 days') ) SELECT U.BlogName, U.PostID, U.LatestNoteTimestamp, U.NotesGatheredDateTime, U.CNT FROM Unioned U WHERE (U.NotesGatheredDateTime < 1786134037 OR U.NotesGatheredDateTime IS NULL) ORDER BY U.NotesGatheredDateTime, U.PostDate DESC, U.BlogName, U.PostID;</sql><sql name="Del Posts">delete from posts where postid in
|
||||
(
|
||||
'741662499571728384',
|
||||
178892849664,
|
||||
178264721139,
|
||||
177012868749,
|
||||
169950081964,
|
||||
755440787056099328
|
||||
)</sql><sql name="notes NO post*">select *
|
||||
-- delete
|
||||
from notes
|
||||
where postid not in (select distinct postid from posts where IsActive = 1)</sql><sql name="SQL 9">SELECT
|
||||
*
|
||||
FROM
|
||||
POSTS P
|
||||
WHERE
|
||||
P.ByLikes = 1
|
||||
AND
|
||||
P.DateCreated > '2026-05-26 17:47:32'
|
||||
ORDER BY
|
||||
P.DateCreated desc</sql><sql name="SQL 13">update posts set IsActive = 0 where blogname IN ( 'shoebiedoo', 'redheaded-girlygirl', 'xlittle-ghost' )</sql><sql name="SQL 14*">update Posts␍
|
||||
set IsActive = 0␍
|
||||
where postid in␍
|
||||
(␍
|
||||
␍
|
||||
␍
|
||||
'731937314675310592'␍
|
||||
␍
|
||||
␍
|
||||
␍
|
||||
)␍
|
||||
␍
|
||||
</sql><current_tab id="7"/></tab_sql></sqlb_project>
|
||||
<?xml version="1.0" encoding="UTF-8"?><sqlb_project><db path="D:/NextCloud/C#/URLNotesGrabberCORE/URLNotesGrabberCORE/TL.db" readonly="0" foreign_keys="1" case_sensitive_like="0" temp_store="0" wal_autocheckpoint="1000" synchronous="2"/><attached/><window><main_tabs open="structure browser pragmas query" current="3"/></window><tab_structure><column_width id="0" width="300"/><column_width id="1" width="0"/><column_width id="2" width="100"/><column_width id="3" width="4486"/><column_width id="4" width="0"/><expanded_item id="0" parent="1"/><expanded_item id="1" parent="1"/><expanded_item id="2" parent="1"/><expanded_item id="3" parent="1"/></tab_structure><tab_browse><table title="Posts" custom_title="0" dock_id="4" table="4,5:mainPosts"/><dock_state state="000000ff00000000fd0000000100000002000005470000029afc0100000006fb000000160064006f0063006b00420072006f00770073006500310100000000000004a10000000000000000fb000000160064006f0063006b00420072006f00770073006500320100000000000004a10000000000000000fb000000160064006f0063006b00420072006f00770073006500330100000000000004a10000000000000000fb000000160064006f0063006b00420072006f00770073006500350100000000000005f40000000000000000fb000000160064006f0063006b00420072006f00770073006500340100000000000005470000011100fffffffb000000160064006f0063006b00420072006f00770073006500340100000000000005f40000000000000000000005470000000000000004000000040000000800000008fc00000000"/><default_encoding codec=""/><browse_table_settings><table schema="main" name="ApiKeyPoolMeta" show_row_id="0" encoding="" plot_x_axis="" unlock_view_pk="_rowid_" freeze_columns="0"><sort/><column_widths><column index="1" value="29"/><column index="2" value="64"/></column_widths><filter_values/><conditional_formats/><row_id_formats/><display_formats/><hidden_columns/><plot_y_axes/><global_filter/></table><table schema="main" name="Blogs" show_row_id="0" encoding="" plot_x_axis="" unlock_view_pk="_rowid_" freeze_columns="0"><sort><column index="7" mode="1"/></sort><column_widths><column index="1" value="257"/><column index="2" value="108"/><column index="3" value="63"/><column index="4" value="156"/><column index="5" value="60"/><column index="6" value="81"/><column index="7" value="85"/><column index="8" value="156"/><column index="9" value="156"/><column index="10" value="151"/><column index="11" value="125"/><column index="12" value="129"/></column_widths><filter_values><column index="4" value="1"/><column index="7" value=">2026-05-27 17:22:36"/></filter_values><conditional_formats/><row_id_formats/><display_formats/><hidden_columns/><plot_y_axes/><global_filter/></table><table schema="main" name="Notes" show_row_id="0" encoding="" plot_x_axis="" unlock_view_pk="_rowid_" freeze_columns="0"><sort/><column_widths><column index="1" value="198"/><column index="2" value="144"/><column index="3" value="251"/><column index="4" value="84"/><column index="5" value="53"/><column index="6" value="300"/><column index="7" value="116"/><column index="8" value="152"/><column index="9" value="89"/><column index="10" value="63"/></column_widths><filter_values/><conditional_formats/><row_id_formats/><display_formats/><hidden_columns/><plot_y_axes/><global_filter/></table><table schema="main" name="Posts" show_row_id="0" encoding="" plot_x_axis="" unlock_view_pk="_rowid_" freeze_columns="0"><sort><column index="14" mode="1"/></sort><column_widths><column index="1" value="236"/><column index="2" value="144"/><column index="3" value="126"/><column index="4" value="32"/><column index="5" value="32"/><column index="6" value="32"/><column index="7" value="32"/><column index="8" value="32"/><column index="9" value="0"/><column index="10" value="0"/><column index="11" value="0"/><column index="12" value="243"/><column index="13" value="300"/><column index="14" value="53"/><column index="15" value="37351"/><column index="16" value="300"/><column index="17" value="41"/><column index="18" value="75"/><column index="19" value="96"/><column index="20" value="300"/><column index="21" value="96"/><column index="22" value="300"/><column index="23" value="300"/><column index="24" value="548"/><column index="25" value="60"/><column index="26" value="213"/><column index="27" value="532"/><column index="28" value="152"/><column index="29" value="152"/><column index="30" value="69"/><column index="31" value="63"/></column_widths><filter_values><column index="2" value="0"/><column index="1" value="734568371821084672"/></filter_values><conditional_formats/><row_id_formats/><display_formats/><hidden_columns><column index="9" value="1"/><column index="10" value="1"/><column index="11" value="1"/></hidden_columns><plot_y_axes/><global_filter/></table></browse_table_settings></tab_browse><tab_sql><sql name="SQL 1">UPDATE Posts
|
||||
SET HasNotesGathered = 0
|
||||
WHERE (BlogName, PostID) IN (
|
||||
SELECT p.BlogName, p.PostID
|
||||
FROM Posts p
|
||||
WHERE p.HasNotesGathered = 1
|
||||
AND P.notesGatheredDatetime < 1774294520
|
||||
AND EXISTS (
|
||||
SELECT 1
|
||||
FROM Notes n
|
||||
WHERE n.PostID = p.PostID
|
||||
AND n.RootBlogId = (SELECT BlogId FROM BlogNames WHERE BlogName = p.BlogName)
|
||||
--AND n.TypeId NOT IN (SELECT TypeId FROM NoteTypes WHERE Type IN ('reblog', 'reply'))
|
||||
)
|
||||
ORDER BY P.PostDate ASC
|
||||
--LIMIT 500
|
||||
);</sql><sql name="Mark Blogs">select *
|
||||
from Blogs
|
||||
--update blogs set HasBeenOutput = 1
|
||||
where HasBeenOutput = 0
|
||||
AND
|
||||
blogname in
|
||||
(
|
||||
'teaberrybee',
|
||||
'reddevilgoddesstoo',
|
||||
'waywardog13',
|
||||
'wzjustbrowsing-blog',
|
||||
'lewerta',
|
||||
'nudenymph',
|
||||
'caylachief'
|
||||
|
||||
)</sql><sql name="New Notes">select P.slug, N.replyText, rbn.BlogName as RootBlogName, n.PostID, nbn.BlogName || '.tumblr.com' as NoteBlogName, DatetimeCrawled, TimeStamp, nt.Type, rbn.BlogName || '.tumblr.com/post/' || n.postid, datetime(timestamp, 'unixepoch')
|
||||
from Notes N
|
||||
inner join Posts P on p.PostID = n.PostID
|
||||
inner join BlogNames rbn on rbn.BlogId = n.RootBlogId
|
||||
inner join BlogNames nbn on nbn.BlogId = n.NoteBlogId
|
||||
inner join NoteTypes nt on nt.TypeId = n.TypeId
|
||||
where
|
||||
DatetimeCrawled > '2026-08-07 11:47:22' and nt.Type like 'r%'
|
||||
and P.IsActive = 1
|
||||
order by n.DatetimeCrawled</sql><sql name="Pull Blogs">SELECT distinct
|
||||
'''' || blogname || ''',',
|
||||
blogs.*
|
||||
, blogname || '.tumblr.com'
|
||||
FROM
|
||||
Blogs
|
||||
inner JOIN
|
||||
Notes on notes.noteBlogId = blogs.BlogId
|
||||
inner JOIN
|
||||
NoteTypes on NoteTypes.TypeId = Notes.TypeId
|
||||
WHERE
|
||||
HasBeenOutput = 0 and NoteTypes.Type = 'reblog'
|
||||
order by
|
||||
NoteTypes.Type desc,
|
||||
DateAdded desc
|
||||
LIMIT 100;</sql><sql name="SQL 7">WITH ReplyCounts AS (
|
||||
SELECT
|
||||
NoteBlogId,
|
||||
COUNT(DISTINCT replyText) AS DistinctReplyCount
|
||||
FROM Notes
|
||||
where replyText <> '.'
|
||||
GROUP BY NoteBlogId
|
||||
)
|
||||
SELECT
|
||||
rbn.BlogName || '.tumblr.com/post/' || n.PostID AS PostURL, postid,
|
||||
nbn.BlogName AS NoteBlogName,
|
||||
n.replyText,
|
||||
c.DistinctReplyCount
|
||||
FROM Notes n
|
||||
JOIN ReplyCounts c ON n.NoteBlogId = c.NoteBlogId
|
||||
JOIN BlogNames rbn ON rbn.BlogId = n.RootBlogId
|
||||
JOIN BlogNames nbn ON nbn.BlogId = n.NoteBlogId
|
||||
JOIN NoteTypes t ON t.TypeId = n.TypeId
|
||||
where replyText <> '.' and t.Type <> 'reply'
|
||||
--AND nbn.BlogName NOT IN ( 'roadblocker21', 'thesaddemon666', 'edwardabbeyhoffman', 'tattedsoldier20', 'zomb-eh', 'animalistic13', 'indken', 'maccloud1592',
|
||||
--'moss-wizard', 'supertrucker12682', 'exploringthrupics', 'padeyepete' )
|
||||
order by c.DistinctReplyCount desc, nbn.BlogName, n.DateModified desc, replyText, rbn.BlogName, PostID</sql><sql name="Collect">WITH PostsWithCount AS ( SELECT P.BlogName, P.PostID, 1925013599 AS LatestNoteTimestamp, P.NotesGatheredDateTime, COUNT(P.PostID) OVER(PARTITION BY P.BlogName) AS CNT, P.HasNotesGathered, P.NotFound, P.PostDate FROM Posts P WHERE COALESCE(P.IsActive, 1) = 1 ), Unioned AS ( SELECT BlogName, PostID, LatestNoteTimestamp, NotesGatheredDateTime, CNT, PostDate FROM PostsWithCount WHERE NotFound = 0 AND HasNotesGathered = 0 UNION SELECT BlogName, PostID, LatestNoteTimestamp, NotesGatheredDateTime, CNT, PostDate FROM PostsWithCount WHERE BlogName = 'zomb-eh' AND NotFound = 0 AND NotesGatheredDateTime < unixepoch('now', 'localtime', '-3 days') ) SELECT U.BlogName, U.PostID, U.LatestNoteTimestamp, U.NotesGatheredDateTime, U.CNT FROM Unioned U WHERE (U.NotesGatheredDateTime < 1786134037 OR U.NotesGatheredDateTime IS NULL) ORDER BY U.NotesGatheredDateTime, U.PostDate DESC, U.BlogName, U.PostID;</sql><sql name="Del Posts">delete from posts where postid in
|
||||
(
|
||||
'741662499571728384',
|
||||
178892849664,
|
||||
178264721139,
|
||||
177012868749,
|
||||
169950081964,
|
||||
755440787056099328
|
||||
)</sql><sql name="notes NO post*">select *
|
||||
-- delete
|
||||
from notes
|
||||
where postid not in (select distinct postid from posts where IsActive = 1)</sql><sql name="SQL 9">SELECT
|
||||
*
|
||||
FROM
|
||||
POSTS P
|
||||
WHERE
|
||||
P.ByLikes = 1
|
||||
AND
|
||||
P.DateCreated > '2026-05-26 17:47:32'
|
||||
ORDER BY
|
||||
P.DateCreated desc</sql><sql name="SQL 13">update posts set IsActive = 0 where blogname IN ( 'shoebiedoo', 'redheaded-girlygirl', 'xlittle-ghost' )</sql><sql name="SQL 14*">update Posts
|
||||
set IsActive = 0
|
||||
where postid in
|
||||
(
|
||||
|
||||
|
||||
'731937314675310592'
|
||||
|
||||
|
||||
|
||||
)
|
||||
|
||||
</sql><current_tab id="7"/></tab_sql></sqlb_project>
|
||||
|
||||
@@ -1,24 +1,30 @@
|
||||
<?xml version="1.0" encoding="UTF-8"?><sqlb_project><db path="" readonly="1" foreign_keys="" case_sensitive_like="" temp_store="" wal_autocheckpoint="" synchronous=""/><attached/><window><main_tabs open="structure browser pragmas query" current="3"/></window><tab_structure><column_width id="0" width="300"/><column_width id="1" width="0"/><column_width id="2" width="100"/><column_width id="3" width="3571"/><column_width id="4" width="0"/></tab_structure><tab_browse><table title="." custom_title="0" dock_id="4" table="0,0:"/><dock_state state="000000ff00000000fd0000000100000002000005f40000030ffc0100000002fb000000160064006f0063006b00420072006f00770073006500310100000000000005f40000000000000000fb000000160064006f0063006b00420072006f00770073006500340100000000ffffffff0000011700ffffff000005f40000000000000004000000040000000800000008fc00000000"/><default_encoding codec=""/><browse_table_settings/></tab_browse><tab_sql><sql name="SQL 1">SELECT
|
||||
BlogName || '.tumblr.com/post/' || postID,
|
||||
|
||||
datetime(NotesGatheredDateTime, 'unixepoch'), *
|
||||
FROM
|
||||
Posts
|
||||
WHERE
|
||||
NotesGatheredDateTime <> 0
|
||||
ORDER BY
|
||||
postdate desc</sql><sql name="SQL 2*">SELECT
|
||||
datetime(TimeStamp, 'unixepoch'),
|
||||
RootBlogName || '.tumblr.com/post/' || N.postid,
|
||||
*,
|
||||
NoteBlogName || '.tumblr.com'
|
||||
FROM
|
||||
Notes N␍
|
||||
inner JOIN␍
|
||||
Posts P on P.PostID = N.PostID and P.BlogName = N.RootBlogName
|
||||
WHERE RootBlogName NOT IN ('xlittle-ghost', 'glimmerin-darlin', 'vvenus-child')␍
|
||||
and type like 'r%'␍
|
||||
and RootBlogName = 'zomb-eh'␍
|
||||
and P.HasImage = 1
|
||||
ORDER BY
|
||||
TimeStamp desc</sql><current_tab id="1"/></tab_sql></sqlb_project>
|
||||
<?xml version="1.0" encoding="UTF-8"?><sqlb_project><db path="" readonly="1" foreign_keys="" case_sensitive_like="" temp_store="" wal_autocheckpoint="" synchronous=""/><attached/><window><main_tabs open="structure browser pragmas query" current="3"/></window><tab_structure><column_width id="0" width="300"/><column_width id="1" width="0"/><column_width id="2" width="100"/><column_width id="3" width="3571"/><column_width id="4" width="0"/></tab_structure><tab_browse><table title="." custom_title="0" dock_id="4" table="0,0:"/><dock_state state="000000ff00000000fd0000000100000002000005f40000030ffc0100000002fb000000160064006f0063006b00420072006f00770073006500310100000000000005f40000000000000000fb000000160064006f0063006b00420072006f00770073006500340100000000ffffffff0000011700ffffff000005f40000000000000004000000040000000800000008fc00000000"/><default_encoding codec=""/><browse_table_settings/></tab_browse><tab_sql><sql name="SQL 1">SELECT
|
||||
BlogName || '.tumblr.com/post/' || postID,
|
||||
|
||||
datetime(NotesGatheredDateTime, 'unixepoch'), *
|
||||
FROM
|
||||
Posts
|
||||
WHERE
|
||||
NotesGatheredDateTime <> 0
|
||||
ORDER BY
|
||||
postdate desc</sql><sql name="SQL 2*">SELECT
|
||||
datetime(TimeStamp, 'unixepoch'),
|
||||
rbn.BlogName || '.tumblr.com/post/' || N.postid,
|
||||
*,
|
||||
nbn.BlogName || '.tumblr.com'
|
||||
FROM
|
||||
Notes N
|
||||
inner JOIN
|
||||
BlogNames rbn on rbn.BlogId = N.RootBlogId
|
||||
inner JOIN
|
||||
BlogNames nbn on nbn.BlogId = N.NoteBlogId
|
||||
inner JOIN
|
||||
NoteTypes t on t.TypeId = N.TypeId
|
||||
inner JOIN
|
||||
Posts P on P.PostID = N.PostID and P.BlogName = rbn.BlogName
|
||||
WHERE rbn.BlogName NOT IN ('xlittle-ghost', 'glimmerin-darlin', 'vvenus-child')
|
||||
and t.Type like 'r%'
|
||||
and rbn.BlogName = 'zomb-eh'
|
||||
and P.HasImage = 1
|
||||
ORDER BY
|
||||
TimeStamp desc</sql><current_tab id="1"/></tab_sql></sqlb_project>
|
||||
|
||||
@@ -213,6 +213,59 @@ namespace URLNotesGrabberCORE
|
||||
|
||||
#endregion IsActive
|
||||
|
||||
#region Notes integer schema
|
||||
|
||||
// Notes stopped storing names on 2026-08-07: RootBlogName/NoteBlogName/Type became
|
||||
// RootBlogId/NoteBlogId/TypeId, resolved through BlogNames and NoteTypes. There is no
|
||||
// compatibility view -- a query naming an old column fails outright, so this is a hard
|
||||
// cut rather than an optional column like IsActive. See TL.db.md.
|
||||
//
|
||||
// Two shapes recur below and are spelled out inline rather than hidden behind a helper,
|
||||
// so that every statement reads as the SQL it actually runs:
|
||||
// (SELECT BlogId FROM BlogNames WHERE BlogName = @name) -- unique-index probe, 20k rows
|
||||
// (SELECT TypeId FROM NoteTypes WHERE Type = 'reply') -- 5 rows, effectively free
|
||||
// Joining Notes to Blogs is the one case that must NOT route through BlogNames: Blogs
|
||||
// carries its own BlogId, so N.NoteBlogId = B.BlogId is a single integer hop. Joining
|
||||
// Notes to Posts is the opposite case -- Posts has only BlogName, so it has to go
|
||||
// through BlogNames.
|
||||
|
||||
/// <summary>
|
||||
/// True when the exception is a duplicate-key collision on Notes. The message embeds the
|
||||
/// primary key's column names, which the integer migration renamed, so this matches on the
|
||||
/// constraint and the table instead of on an exact column list -- a literal comparison
|
||||
/// silently inverts into "log every error" the next time a column is renamed.
|
||||
/// </summary>
|
||||
private static bool IsNotesDuplicateKey(Exception ex)
|
||||
{
|
||||
return ex.Message.Contains("UNIQUE constraint failed", StringComparison.OrdinalIgnoreCase)
|
||||
&& ex.Message.Contains("Notes.", StringComparison.OrdinalIgnoreCase);
|
||||
}
|
||||
|
||||
/// <summary>
|
||||
/// Gives a blog name an ID if it does not have one. No read-back and no round trip -- a
|
||||
/// name that is already registered keeps the ID that 1.18M Notes rows point at.
|
||||
/// </summary>
|
||||
private static void RegisterBlogName(SQLiteConnection connection, SQLiteTransaction? transaction, string blogName)
|
||||
{
|
||||
using SQLiteCommand command = new SQLiteCommand("INSERT OR IGNORE INTO BlogNames (BlogName) VALUES (@BlogName)", connection, transaction);
|
||||
command.Parameters.AddWithValue("@BlogName", blogName);
|
||||
command.ExecuteNonQuery();
|
||||
}
|
||||
|
||||
/// <summary>
|
||||
/// Same, for a note type. NoteTypes is a table rather than a CHECK constraint precisely so
|
||||
/// that a type this crawler has not seen before is an INSERT and not a schema migration --
|
||||
/// without this the type would resolve to NULL and fail the NOT NULL on Notes.TypeId.
|
||||
/// </summary>
|
||||
private static void RegisterNoteType(SQLiteConnection connection, SQLiteTransaction? transaction, string type)
|
||||
{
|
||||
using SQLiteCommand command = new SQLiteCommand("INSERT OR IGNORE INTO NoteTypes (Type) VALUES (@Type)", connection, transaction);
|
||||
command.Parameters.AddWithValue("@Type", type);
|
||||
command.ExecuteNonQuery();
|
||||
}
|
||||
|
||||
#endregion Notes integer schema
|
||||
|
||||
public static string Q(string input)
|
||||
{
|
||||
return "'" + input.Replace("'", "''") + "'";
|
||||
@@ -371,7 +424,10 @@ namespace URLNotesGrabberCORE
|
||||
connection.Close();
|
||||
connection.Open();
|
||||
|
||||
string addColumnSql = "ALTER TABLE Notes ADD COLUMN replyText TEXT DEFAULT '.';";
|
||||
// No column default: the migrated schema dropped the DEFAULT '.' that
|
||||
// is how 1.1M rows acquired a placeholder nobody wrote. New rows get
|
||||
// NULL, which every reader here already treats as "no reply text".
|
||||
string addColumnSql = "ALTER TABLE Notes ADD COLUMN replyText TEXT;";
|
||||
using (SQLiteCommand addCommand = new SQLiteCommand(addColumnSql, connection))
|
||||
{
|
||||
addCommand.ExecuteNonQuery();
|
||||
@@ -534,12 +590,16 @@ namespace URLNotesGrabberCORE
|
||||
}
|
||||
}
|
||||
|
||||
public static void AddPost(string blogName, long postID, string reblogURL, string postDate, string postURL, string slug, string reblogKey, string reblogName, string summary, string quote, string body, string tags, string link, string photoURL, string photoCaption, string downloadedFiles, string audioCaption, string question, string answer, string title, bool hasImage, bool byLikes = false, string? DBPath = null, string? rootBlogName = null, string? rootURL = null)
|
||||
// postType: a canonical PostTypes name, or null when the caller has no trustworthy type.
|
||||
// Null is stored as NULL rather than guessed at -- OutputMode skips untyped rows, so a
|
||||
// null costs one export line, whereas a wrong value would create a wrongly named file.
|
||||
public static void AddPost(string blogName, long postID, string reblogURL, string postDate, string postURL, string slug, string reblogKey, string reblogName, string summary, string quote, string body, string tags, string link, string photoURL, string photoCaption, string downloadedFiles, string audioCaption, string question, string answer, string title, bool hasImage, bool byLikes = false, string? DBPath = null, string? rootBlogName = null, string? rootURL = null, string? postType = null)
|
||||
{
|
||||
DBPath ??= GetDefaultDbPath();
|
||||
postType = PostTypes.Normalize(postType);
|
||||
try { AddBlog(blogName, byLikes, DBPath); } catch { }
|
||||
try { UpdatePostSetDate(blogName, postID, postDate, DBPath); } catch { }
|
||||
try { UpdatePost(blogName, postID, reblogURL, postDate, postURL, slug, reblogKey, reblogName, summary, quote, body, tags, link, photoURL, photoCaption, downloadedFiles, audioCaption, question, answer, title, hasImage, byLikes, DBPath, rootBlogName, rootURL); } catch { }
|
||||
try { UpdatePost(blogName, postID, reblogURL, postDate, postURL, slug, reblogKey, reblogName, summary, quote, body, tags, link, photoURL, photoCaption, downloadedFiles, audioCaption, question, answer, title, hasImage, byLikes, DBPath, rootBlogName, rootURL, postType); } catch { }
|
||||
|
||||
SQLiteConnection connection;
|
||||
bool ownsConnection;
|
||||
@@ -581,9 +641,10 @@ namespace URLNotesGrabberCORE
|
||||
RootBlogName,
|
||||
RootURL,
|
||||
HasImage,
|
||||
ByLikes
|
||||
ByLikes,
|
||||
PostType
|
||||
) VALUES (" +
|
||||
Q(blogName) + ", " + postID + ", " + Q(reblogURL) + ", " + Q(postDate) + ", " + Q(postURL) + ", " + Q(slug) + ", " + Q(reblogKey) + ", " + Q(reblogName) + ", " + Q(summary) + ", " + Q(quote) + ", " + Q(body) + ", " + Q(tags) + ", " + Q(link) + ", " + Q(photoURL) + ", " + Q(photoCaption) + ", " + Q(downloadedFiles) + ", " + Q(audioCaption) + ", " + Q(question) + ", " + Q(answer) + ", " + Q(title) + ", " + Q(DateTime.Now.ToString("yyyy-MM-dd HH:mm:ss")) + ", " + Q(DateTime.Now.ToString("yyyy-MM-dd HH:mm:ss")) + ", " + Q(rootBlogName ?? ".") + ", " + Q(rootURL ?? ".") + ", " + (hasImage ? 1 : 0) + ", " + (byLikes ? 1 : 0) + ")";
|
||||
Q(blogName) + ", " + postID + ", " + Q(reblogURL) + ", " + Q(postDate) + ", " + Q(postURL) + ", " + Q(slug) + ", " + Q(reblogKey) + ", " + Q(reblogName) + ", " + Q(summary) + ", " + Q(quote) + ", " + Q(body) + ", " + Q(tags) + ", " + Q(link) + ", " + Q(photoURL) + ", " + Q(photoCaption) + ", " + Q(downloadedFiles) + ", " + Q(audioCaption) + ", " + Q(question) + ", " + Q(answer) + ", " + Q(title) + ", " + Q(DateTime.Now.ToString("yyyy-MM-dd HH:mm:ss")) + ", " + Q(DateTime.Now.ToString("yyyy-MM-dd HH:mm:ss")) + ", " + Q(rootBlogName ?? ".") + ", " + Q(rootURL ?? ".") + ", " + (hasImage ? 1 : 0) + ", " + (byLikes ? 1 : 0) + ", " + (postType == null ? "NULL" : Q(postType)) + ")";
|
||||
SQLiteCommand command = new SQLiteCommand(sql, connection);
|
||||
|
||||
int rowsInserted = 0;
|
||||
@@ -703,63 +764,85 @@ namespace URLNotesGrabberCORE
|
||||
try { AddBlog(noteBlogName, false, DBPath); } catch { }
|
||||
|
||||
using SQLiteConnection connection2 = new SQLiteConnection("Data Source=" + DBPath);
|
||||
int rowsInserted = 0;
|
||||
|
||||
try
|
||||
{
|
||||
connection2.Open();
|
||||
|
||||
// INSERT OR IGNORE, and no IsActive in the column list: re-crawling a note
|
||||
// that was removed elsewhere leaves the existing row -- and its flag -- alone.
|
||||
string sql = "INSERT OR IGNORE INTO Notes (rootBlogName, noteBlogName, PostID, TimeStamp, Type, DatetimeCrawled, DateModified, DateCreated) values(@rootBlogName, @noteBlogName, @PostID, @TimeStamp, @Type, @DatetimeCrawled, @DateModified, @DateCreated)";
|
||||
using (SQLiteCommand command = new SQLiteCommand(sql, connection2))
|
||||
// Notes stores integer IDs, so both participants and the type have to exist in
|
||||
// their lookup table before the note can point at them.
|
||||
//
|
||||
// All four statements run in one transaction so a crash cannot leave a name or a
|
||||
// type registered with no note. The transaction is committed before the console
|
||||
// output below, which sleeps -- a write lock must not be held across that.
|
||||
using (SQLiteTransaction transaction = connection2.BeginTransaction())
|
||||
{
|
||||
command.Parameters.AddWithValue("@rootBlogName", rootBlogName);
|
||||
command.Parameters.AddWithValue("@noteBlogName", noteBlogName);
|
||||
command.Parameters.AddWithValue("@PostID", postID);
|
||||
command.Parameters.AddWithValue("@TimeStamp", timestamp);
|
||||
command.Parameters.AddWithValue("@Type", type ?? string.Empty);
|
||||
command.Parameters.AddWithValue("@DatetimeCrawled", DateTime.Now.ToString("yyyy-MM-dd HH:mm:ss"));
|
||||
command.Parameters.AddWithValue("@DateModified", DateTime.Now.ToString("yyyy-MM-dd HH:mm:ss"));
|
||||
command.Parameters.AddWithValue("@DateCreated", DateTime.Now.ToString("yyyy-MM-dd HH:mm:ss"));
|
||||
RegisterBlogName(connection2, transaction, rootBlogName);
|
||||
RegisterBlogName(connection2, transaction, noteBlogName);
|
||||
RegisterNoteType(connection2, transaction, type ?? string.Empty);
|
||||
|
||||
int rowsInserted = command.ExecuteNonQuery();
|
||||
// INSERT OR IGNORE, and no IsActive in the column list: re-crawling a note
|
||||
// that was removed elsewhere leaves the existing row -- and its flag -- alone.
|
||||
string sql = "INSERT OR IGNORE INTO Notes (RootBlogId, NoteBlogId, PostID, TimeStamp, TypeId, DatetimeCrawled, DateModified, DateCreated) " +
|
||||
"SELECT (SELECT BlogId FROM BlogNames WHERE BlogName = @rootBlogName), " +
|
||||
" (SELECT BlogId FROM BlogNames WHERE BlogName = @noteBlogName), " +
|
||||
" @PostID, @TimeStamp, " +
|
||||
" (SELECT TypeId FROM NoteTypes WHERE Type = @Type), " +
|
||||
" @DatetimeCrawled, @DateModified, @DateCreated";
|
||||
|
||||
if (rowsInserted == 1)
|
||||
using (SQLiteCommand command = new SQLiteCommand(sql, connection2, transaction))
|
||||
{
|
||||
ConsoleColor previousColor = Console.ForegroundColor;
|
||||
Console.ForegroundColor = ConsoleColor.Green;
|
||||
Console.WriteLine("{2}\t{0}\t{1}", UnixTimeStampToDateTime(timestamp), noteBlogName, type);
|
||||
Console.ForegroundColor = previousColor;
|
||||
Thread.Sleep(250); // Brief pause to make new notes more noticeable in the console output
|
||||
}
|
||||
else
|
||||
{
|
||||
Console.WriteLine("{2}\t{0}\t{1}", UnixTimeStampToDateTime(timestamp), noteBlogName, type);
|
||||
command.Parameters.AddWithValue("@rootBlogName", rootBlogName);
|
||||
command.Parameters.AddWithValue("@noteBlogName", noteBlogName);
|
||||
command.Parameters.AddWithValue("@PostID", postID);
|
||||
command.Parameters.AddWithValue("@TimeStamp", timestamp);
|
||||
command.Parameters.AddWithValue("@Type", type ?? string.Empty);
|
||||
command.Parameters.AddWithValue("@DatetimeCrawled", DateTime.Now.ToString("yyyy-MM-dd HH:mm:ss"));
|
||||
command.Parameters.AddWithValue("@DateModified", DateTime.Now.ToString("yyyy-MM-dd HH:mm:ss"));
|
||||
command.Parameters.AddWithValue("@DateCreated", DateTime.Now.ToString("yyyy-MM-dd HH:mm:ss"));
|
||||
|
||||
rowsInserted = command.ExecuteNonQuery();
|
||||
}
|
||||
|
||||
// Only update HasBeenOutput if a new note was inserted
|
||||
if (rowsInserted == 1)
|
||||
transaction.Commit();
|
||||
}
|
||||
|
||||
if (rowsInserted == 1)
|
||||
{
|
||||
ConsoleColor previousColor = Console.ForegroundColor;
|
||||
Console.ForegroundColor = ConsoleColor.Green;
|
||||
Console.WriteLine("{2}\t{0}\t{1}", UnixTimeStampToDateTime(timestamp), noteBlogName, type);
|
||||
Console.ForegroundColor = previousColor;
|
||||
Thread.Sleep(125); // Brief pause to make new notes more noticeable in the console output
|
||||
}
|
||||
else
|
||||
{
|
||||
Console.WriteLine("{2}\t{0}\t{1}", UnixTimeStampToDateTime(timestamp), noteBlogName, type);
|
||||
}
|
||||
|
||||
// Only update HasBeenOutput if a new note was inserted
|
||||
if (rowsInserted == 1)
|
||||
{
|
||||
try
|
||||
{
|
||||
try
|
||||
// HasBeenOutput IS NULL still counts as a change: the selection queries
|
||||
// test HasBeenOutput = 0, which a NULL would never match.
|
||||
string updateSql = "UPDATE Blogs SET HasBeenOutput = 0, DateModified = @DateModified WHERE BlogName = @BlogName AND (HasBeenOutput IS NULL OR HasBeenOutput <> 0)";
|
||||
using (var updateCommand = new SQLiteCommand(updateSql, connection2))
|
||||
{
|
||||
// HasBeenOutput IS NULL still counts as a change: the selection queries
|
||||
// test HasBeenOutput = 0, which a NULL would never match.
|
||||
string updateSql = "UPDATE Blogs SET HasBeenOutput = 0, DateModified = @DateModified WHERE BlogName = @BlogName AND (HasBeenOutput IS NULL OR HasBeenOutput <> 0)";
|
||||
using (var updateCommand = new SQLiteCommand(updateSql, connection2))
|
||||
{
|
||||
updateCommand.Parameters.AddWithValue("@BlogName", noteBlogName);
|
||||
updateCommand.Parameters.AddWithValue("@DateModified", DateTime.Now.ToString("yyyy-MM-dd HH:mm:ss"));
|
||||
updateCommand.ExecuteNonQuery();
|
||||
}
|
||||
updateCommand.Parameters.AddWithValue("@BlogName", noteBlogName);
|
||||
updateCommand.Parameters.AddWithValue("@DateModified", DateTime.Now.ToString("yyyy-MM-dd HH:mm:ss"));
|
||||
updateCommand.ExecuteNonQuery();
|
||||
}
|
||||
catch { }
|
||||
}
|
||||
catch { }
|
||||
}
|
||||
}
|
||||
catch (Exception ex)
|
||||
{
|
||||
// Breakpoint here
|
||||
if (ex.Message != "constraint failed\r\nUNIQUE constraint failed: Notes.RootBlogName, Notes.PostID, Notes.TimeStamp, Notes.Type, Notes.NoteBlogName")
|
||||
if (!IsNotesDuplicateKey(ex))
|
||||
{
|
||||
Console.WriteLine(ex.Message);
|
||||
Console.WriteLine("^^^^^ - SHORTCUT");
|
||||
@@ -777,12 +860,16 @@ namespace URLNotesGrabberCORE
|
||||
/// <param name="withoutNotesOnly"></param>
|
||||
/// <param name="DBPath"></param>
|
||||
/// <returns>blogName, postID, lastNoteTimestamp, notesGatheredTimestamp</returns>
|
||||
public static List<Tuple<string, long, long, long>> GetPosts(bool withoutNotesOnly = false, DateTime? beforeDate = null, string? DBPath = null)
|
||||
public static List<Tuple<string, long, long, long>> GetPosts(bool withoutNotesOnly = false, DateTime? beforeDate = null, string? blogName = null, string? DBPath = null)
|
||||
{
|
||||
DBPath ??= GetDefaultDbPath();
|
||||
using SQLiteConnection connection = new SQLiteConnection("Data Source=" + DBPath);
|
||||
List<Tuple<string, long, long, long>> posts = new List<Tuple<string, long, long, long>>();
|
||||
|
||||
// A blog filter matches BlogName exactly: the column is BINARY-collated and leads the
|
||||
// Posts primary key, so "= @blogName" rides that index instead of scanning 1.18M rows.
|
||||
bool filterByBlog = !string.IsNullOrWhiteSpace(blogName);
|
||||
|
||||
try
|
||||
{
|
||||
connection.Open();
|
||||
@@ -798,31 +885,17 @@ namespace URLNotesGrabberCORE
|
||||
beforeDateFilter = $"WHERE (U.NotesGatheredDateTime < {unixTimestamp} OR U.NotesGatheredDateTime IS NULL)" + Environment.NewLine;
|
||||
}
|
||||
|
||||
sql = "WITH PostsWithCount AS" + Environment.NewLine +
|
||||
"(" + Environment.NewLine +
|
||||
" SELECT " + Environment.NewLine +
|
||||
" P.BlogName," + Environment.NewLine +
|
||||
" P.PostID," + Environment.NewLine +
|
||||
" 1925013599 AS LatestNoteTimestamp," + Environment.NewLine +
|
||||
" P.NotesGatheredDateTime," + Environment.NewLine +
|
||||
" COUNT(P.PostID) OVER(PARTITION BY P.BlogName) AS CNT," + Environment.NewLine +
|
||||
" P.HasNotesGathered," + Environment.NewLine +
|
||||
" P.NotFound," + Environment.NewLine +
|
||||
" P.PostDate" + Environment.NewLine +
|
||||
" FROM Posts P" + WhereIsActive("Posts", "P", DBPath) + Environment.NewLine +
|
||||
")," + Environment.NewLine +
|
||||
"Unioned AS" + Environment.NewLine +
|
||||
"(" + Environment.NewLine +
|
||||
" SELECT " + Environment.NewLine +
|
||||
" BlogName," + Environment.NewLine +
|
||||
" PostID," + Environment.NewLine +
|
||||
" LatestNoteTimestamp," + Environment.NewLine +
|
||||
" NotesGatheredDateTime," + Environment.NewLine +
|
||||
" CNT," + Environment.NewLine +
|
||||
" PostDate" + Environment.NewLine +
|
||||
" FROM PostsWithCount" + Environment.NewLine +
|
||||
" WHERE NotFound = 0" + Environment.NewLine +
|
||||
" AND HasNotesGathered = 0" + Environment.NewLine +
|
||||
// Same WHERE/AND juggling WhereIsActive does, extended to the optional blog predicate:
|
||||
// either clause may be absent, so the first one present has to open the WHERE.
|
||||
string sourceClause = AndIsActive("Posts", "P", DBPath) + (filterByBlog ? " AND P.BlogName = @blogName" : string.Empty);
|
||||
string sourceFilter = sourceClause.Length == 0 ? string.Empty : " WHERE" + sourceClause.Substring(" AND".Length);
|
||||
|
||||
// The zomb-eh branch re-queues that blog's *already collected* posts every 3 days. It needs
|
||||
// no blog-filter handling of its own: it reads PostsWithCount, which the filter has already
|
||||
// scoped, so it contributes its rows when the filter names zomb-eh and nothing otherwise.
|
||||
// That keeps a filtered worklist a strict subset of the unfiltered one -- "--collect 1 X"
|
||||
// returns exactly the rows "--collect 1" would have returned for X.
|
||||
string refreshBranch =
|
||||
"" + Environment.NewLine +
|
||||
" UNION " + Environment.NewLine +
|
||||
"" + Environment.NewLine +
|
||||
@@ -836,7 +909,34 @@ namespace URLNotesGrabberCORE
|
||||
" FROM PostsWithCount" + Environment.NewLine +
|
||||
" WHERE BlogName = 'zomb-eh'" + Environment.NewLine +
|
||||
" AND NotFound = 0" + Environment.NewLine +
|
||||
" AND NotesGatheredDateTime < unixepoch('now', 'localtime', '-3 days')" + Environment.NewLine +
|
||||
" AND NotesGatheredDateTime < unixepoch('now', 'localtime', '-3 days')" + Environment.NewLine;
|
||||
|
||||
sql = "WITH PostsWithCount AS" + Environment.NewLine +
|
||||
"(" + Environment.NewLine +
|
||||
" SELECT " + Environment.NewLine +
|
||||
" P.BlogName," + Environment.NewLine +
|
||||
" P.PostID," + Environment.NewLine +
|
||||
" 1925013599 AS LatestNoteTimestamp," + Environment.NewLine +
|
||||
" P.NotesGatheredDateTime," + Environment.NewLine +
|
||||
" COUNT(P.PostID) OVER(PARTITION BY P.BlogName) AS CNT," + Environment.NewLine +
|
||||
" P.HasNotesGathered," + Environment.NewLine +
|
||||
" P.NotFound," + Environment.NewLine +
|
||||
" P.PostDate" + Environment.NewLine +
|
||||
" FROM Posts P" + sourceFilter + Environment.NewLine +
|
||||
")," + Environment.NewLine +
|
||||
"Unioned AS" + Environment.NewLine +
|
||||
"(" + Environment.NewLine +
|
||||
" SELECT " + Environment.NewLine +
|
||||
" BlogName," + Environment.NewLine +
|
||||
" PostID," + Environment.NewLine +
|
||||
" LatestNoteTimestamp," + Environment.NewLine +
|
||||
" NotesGatheredDateTime," + Environment.NewLine +
|
||||
" CNT," + Environment.NewLine +
|
||||
" PostDate" + Environment.NewLine +
|
||||
" FROM PostsWithCount" + Environment.NewLine +
|
||||
" WHERE NotFound = 0" + Environment.NewLine +
|
||||
" AND HasNotesGathered = 0" + Environment.NewLine +
|
||||
refreshBranch +
|
||||
")" + Environment.NewLine +
|
||||
"SELECT" + Environment.NewLine +
|
||||
" U.BlogName," + Environment.NewLine +
|
||||
@@ -850,6 +950,10 @@ namespace URLNotesGrabberCORE
|
||||
}
|
||||
else
|
||||
{
|
||||
// The LEFT OUTER JOIN to Notes that used to sit here has been dropped rather
|
||||
// than ported. Nothing was selected from it, a LEFT JOIN cannot remove a row,
|
||||
// and the GROUP BY below collapsed the rows it duplicated -- so it could not
|
||||
// affect the result, and it cost a join against 1.18M rows on every pass.
|
||||
sql = "SELECT " +
|
||||
" MAX(Posts.BlogName) as BlogName, " + Environment.NewLine +
|
||||
" Posts.PostID, " + Environment.NewLine +
|
||||
@@ -859,11 +963,14 @@ namespace URLNotesGrabberCORE
|
||||
"FROM " + Environment.NewLine +
|
||||
" Posts " + Environment.NewLine +
|
||||
" LEFT OUTER JOIN " + Environment.NewLine +
|
||||
" Notes ON Notes.RootBlogName = Posts.BlogName AND Notes.PostID = Posts.PostID " + Environment.NewLine +
|
||||
" LEFT OUTER JOIN " + Environment.NewLine +
|
||||
" ( select BlogName, count(PostID) as CNT from Posts" + WhereIsActive("Posts", "", DBPath) + " group by BlogName) CNT on CNT.blogName = Posts.BlogName " +
|
||||
"WHERE NotFound = 0 " + AndIsActive("Posts", "Posts", DBPath) + Environment.NewLine;
|
||||
|
||||
if (filterByBlog)
|
||||
{
|
||||
sql += " AND Posts.BlogName = @blogName " + Environment.NewLine;
|
||||
}
|
||||
|
||||
if (beforeDate.HasValue)
|
||||
{
|
||||
long unixTimestamp = new DateTimeOffset(beforeDate.Value).ToUnixTimeSeconds();
|
||||
@@ -883,6 +990,9 @@ namespace URLNotesGrabberCORE
|
||||
|
||||
using (SQLiteCommand command = new SQLiteCommand(sql, connection))
|
||||
{
|
||||
if (filterByBlog)
|
||||
command.Parameters.AddWithValue("@blogName", blogName);
|
||||
|
||||
using (SQLiteDataReader reader = command.ExecuteReader())
|
||||
{
|
||||
while (reader.Read())
|
||||
@@ -971,7 +1081,11 @@ namespace URLNotesGrabberCORE
|
||||
{
|
||||
connection.Open();
|
||||
|
||||
string sql = "SELECT distinct RootBlogName as blogName, postID FROM Notes WHERE Notes.type = 'reply'" + AndIsActive("Notes", "Notes", DBPath) + " order by RootBlogName, PostID";
|
||||
string sql = "SELECT DISTINCT BN.BlogName as blogName, N.PostID" +
|
||||
" FROM Notes N" +
|
||||
" INNER JOIN BlogNames BN ON BN.BlogId = N.RootBlogId" +
|
||||
" WHERE N.TypeId = (SELECT TypeId FROM NoteTypes WHERE Type = 'reply')" + AndIsActive("Notes", "N", DBPath) +
|
||||
" ORDER BY BN.BlogName, N.PostID";
|
||||
|
||||
using (SQLiteCommand command = new SQLiteCommand(sql, connection))
|
||||
{
|
||||
@@ -1011,12 +1125,15 @@ namespace URLNotesGrabberCORE
|
||||
{
|
||||
connection.Open();
|
||||
|
||||
string sql = @"SELECT DISTINCT Notes.RootBlogName as blogName, Notes.PostID,
|
||||
MAX(Notes.timestamp) as LatestTimestamp
|
||||
FROM Notes
|
||||
WHERE Notes.type = 'reply'
|
||||
AND (Notes.replyText IS NULL OR Notes.replyText = '' OR Notes.replyText = '.')" + AndIsActive("Notes", "Notes", DBPath) + @"
|
||||
GROUP BY Notes.RootBlogName, Notes.PostID
|
||||
// Grouped on the integer rather than the name: the group key is what gets sorted,
|
||||
// and BN.BlogName comes along for free off the join.
|
||||
string sql = @"SELECT BN.BlogName as blogName, N.PostID,
|
||||
MAX(N.TimeStamp) as LatestTimestamp
|
||||
FROM Notes N
|
||||
INNER JOIN BlogNames BN ON BN.BlogId = N.RootBlogId
|
||||
WHERE N.TypeId = (SELECT TypeId FROM NoteTypes WHERE Type = 'reply')
|
||||
AND (N.replyText IS NULL OR N.replyText = '' OR N.replyText = '.')" + AndIsActive("Notes", "N", DBPath) + @"
|
||||
GROUP BY N.RootBlogId, N.PostID
|
||||
ORDER BY LatestTimestamp ASC
|
||||
LIMIT @limit";
|
||||
|
||||
@@ -1058,11 +1175,15 @@ namespace URLNotesGrabberCORE
|
||||
{
|
||||
connection.Open();
|
||||
|
||||
string sql = @"SELECT DISTINCT P.BlogName, P.PostID, MAX(N.timestamp) as LatestTimestamp
|
||||
// Posts carries only BlogName, so this is the one join to Notes that has to go
|
||||
// through BlogNames -- there is no Posts.BlogId to hop on. The name predicate is
|
||||
// pushed into the 20k-row lookup, which then feeds integers to the Notes key.
|
||||
string sql = @"SELECT DISTINCT P.BlogName, P.PostID, MAX(N.TimeStamp) as LatestTimestamp
|
||||
FROM Posts P
|
||||
INNER JOIN Notes N ON N.PostID = P.PostID AND N.RootBlogName = P.BlogName
|
||||
INNER JOIN BlogNames RBN ON RBN.BlogName = P.BlogName
|
||||
INNER JOIN Notes N ON N.RootBlogId = RBN.BlogId AND N.PostID = P.PostID
|
||||
WHERE P.NotFound = 0
|
||||
AND N.type = 'reply'
|
||||
AND N.TypeId = (SELECT TypeId FROM NoteTypes WHERE Type = 'reply')
|
||||
AND (N.replyText IS NULL OR N.replyText = '' OR N.replyText = '.')" + AndIsActive("Posts", "P", DBPath) + AndIsActive("Notes", "N", DBPath) + @"
|
||||
GROUP BY P.BlogName, P.PostID
|
||||
ORDER BY LatestTimestamp ASC";
|
||||
@@ -1160,6 +1281,11 @@ namespace URLNotesGrabberCORE
|
||||
{
|
||||
connection.Open();
|
||||
string sql;
|
||||
|
||||
// The two Notes branches below join on Blogs.BlogId, which is NULL for the 168k
|
||||
// registry rows that have never appeared in a note. The inner join drops them,
|
||||
// which is correct here -- both branches already require a note to exist -- but
|
||||
// it is the wrong shape for anything that lists the registry.
|
||||
if (!string.IsNullOrEmpty(specificBlog))
|
||||
{
|
||||
// Specific blog: always process, bypass cooldown
|
||||
@@ -1179,12 +1305,12 @@ namespace URLNotesGrabberCORE
|
||||
COALESCE(B.LikesCursor, 0),
|
||||
COALESCE(B.LikesNewestTimestamp, 0)
|
||||
FROM Blogs B
|
||||
INNER JOIN Notes N ON N.NoteBlogName = B.BlogName
|
||||
INNER JOIN Notes N ON N.NoteBlogId = B.BlogId
|
||||
WHERE N.TimeStamp >= 1535778000
|
||||
AND N.rootBlogName = B.BlogName
|
||||
AND N.RootBlogId = B.BlogId
|
||||
AND B.IsActive = 1" + AndIsActive("Notes", "N", DBPath) + @"
|
||||
GROUP BY B.BlogName
|
||||
ORDER BY MIN(N.Timestamp);";
|
||||
ORDER BY MIN(N.TimeStamp);";
|
||||
}
|
||||
else
|
||||
{
|
||||
@@ -1194,9 +1320,9 @@ namespace URLNotesGrabberCORE
|
||||
COALESCE(B.LikesCursor, 0),
|
||||
COALESCE(B.LikesNewestTimestamp, 0)
|
||||
FROM Blogs B
|
||||
INNER JOIN Notes N ON N.NoteBlogName = B.BlogName
|
||||
INNER JOIN Notes N ON N.NoteBlogId = B.BlogId
|
||||
WHERE N.TimeStamp >= 1535778000
|
||||
AND N.rootBlogName = B.BlogName
|
||||
AND N.RootBlogId = B.BlogId
|
||||
AND B.IsActive = 1" + AndIsActive("Notes", "N", DBPath) + @"
|
||||
AND (
|
||||
B.LikesPulled = 0
|
||||
@@ -1204,7 +1330,7 @@ namespace URLNotesGrabberCORE
|
||||
< (CAST(strftime('%s','now') AS INTEGER) - (@cooldownDays * 86400))
|
||||
)
|
||||
GROUP BY B.BlogName
|
||||
ORDER BY MIN(N.Timestamp);";
|
||||
ORDER BY MIN(N.TimeStamp);";
|
||||
}
|
||||
|
||||
using (SQLiteCommand command = new SQLiteCommand(sql, connection))
|
||||
@@ -1244,11 +1370,14 @@ namespace URLNotesGrabberCORE
|
||||
try
|
||||
{
|
||||
connection.Open();
|
||||
// Blogs is reached in one integer hop off Blogs.BlogId, not through BlogNames --
|
||||
// that would add a hop and end in the text comparison the migration removed.
|
||||
// The negated form is only correct because Notes.TypeId is NOT NULL.
|
||||
string sql = "";
|
||||
if (reblogsOnly)
|
||||
sql = "SELECT NoteBlogName as blogName, count(*) FROM notes INNER JOIN blogs ON blogs.BlogName = notes.NoteBlogName WHERE blogs.IsActive = @isActive" + AndIsActive("Notes", "notes", DBPath) + " AND type IN ('reblog', 'reply', 'posted') AND HasBeenOutput = 0 GROUP BY NoteBlogName ORDER BY count(*) DESC, BlogName LIMIT @top";
|
||||
sql = "SELECT B.BlogName as blogName, count(*) FROM Notes N INNER JOIN Blogs B ON B.BlogId = N.NoteBlogId WHERE B.IsActive = @isActive" + AndIsActive("Notes", "N", DBPath) + " AND N.TypeId IN (SELECT TypeId FROM NoteTypes WHERE Type IN ('reblog', 'reply', 'posted')) AND B.HasBeenOutput = 0 GROUP BY N.NoteBlogId ORDER BY count(*) DESC, B.BlogName LIMIT @top";
|
||||
else
|
||||
sql = "SELECT NoteBlogName as blogName, count(*) FROM notes INNER JOIN blogs ON blogs.BlogName = notes.NoteBlogName WHERE blogs.IsActive = @isActive" + AndIsActive("Notes", "notes", DBPath) + " AND type NOT IN ('reblog', 'reply', 'posted') AND HasBeenOutput = 0 GROUP BY NoteBlogName ORDER BY count(*) DESC, BlogName LIMIT @top";
|
||||
sql = "SELECT B.BlogName as blogName, count(*) FROM Notes N INNER JOIN Blogs B ON B.BlogId = N.NoteBlogId WHERE B.IsActive = @isActive" + AndIsActive("Notes", "N", DBPath) + " AND N.TypeId NOT IN (SELECT TypeId FROM NoteTypes WHERE Type IN ('reblog', 'reply', 'posted')) AND B.HasBeenOutput = 0 GROUP BY N.NoteBlogId ORDER BY count(*) DESC, B.BlogName LIMIT @top";
|
||||
|
||||
using (SQLiteCommand command = new SQLiteCommand(sql, connection))
|
||||
{
|
||||
@@ -1287,9 +1416,9 @@ namespace URLNotesGrabberCORE
|
||||
connection.Open();
|
||||
string sql = "";
|
||||
if (reblogsOnly)
|
||||
sql = "SELECT NoteBlogName as blogName, count(*) FROM notes INNER JOIN blogs ON blogs.BlogName = notes.NoteBlogName WHERE blogs.IsActive = @isActive" + AndIsActive("Notes", "notes", DBPath) + " AND type IN ('reblog', 'reply', 'posted') AND HasBeenOutput = 0 GROUP BY NoteBlogName ORDER BY count(*) DESC, BlogName LIMIT @top";
|
||||
sql = "SELECT B.BlogName as blogName, count(*) FROM Notes N INNER JOIN Blogs B ON B.BlogId = N.NoteBlogId WHERE B.IsActive = @isActive" + AndIsActive("Notes", "N", DBPath) + " AND N.TypeId IN (SELECT TypeId FROM NoteTypes WHERE Type IN ('reblog', 'reply', 'posted')) AND B.HasBeenOutput = 0 GROUP BY N.NoteBlogId ORDER BY count(*) DESC, B.BlogName LIMIT @top";
|
||||
else
|
||||
sql = "SELECT NoteBlogName as blogName, count(*) FROM notes INNER JOIN blogs ON blogs.BlogName = notes.NoteBlogName WHERE blogs.IsActive = @isActive" + AndIsActive("Notes", "notes", DBPath) + " AND HasBeenOutput = 0 GROUP BY NoteBlogName ORDER BY count(*) DESC, BlogName LIMIT @top";
|
||||
sql = "SELECT B.BlogName as blogName, count(*) FROM Notes N INNER JOIN Blogs B ON B.BlogId = N.NoteBlogId WHERE B.IsActive = @isActive" + AndIsActive("Notes", "N", DBPath) + " AND B.HasBeenOutput = 0 GROUP BY N.NoteBlogId ORDER BY count(*) DESC, B.BlogName LIMIT @top";
|
||||
|
||||
using (SQLiteCommand command = new SQLiteCommand(sql, connection))
|
||||
{
|
||||
@@ -1528,7 +1657,10 @@ namespace URLNotesGrabberCORE
|
||||
{
|
||||
connection.Open();
|
||||
|
||||
string sql = "UPDATE Notes SET timestamp = @timestamp, DateModified = @dateModified WHERE rootBlogName = @rootBlogName AND noteBlogName = @noteBlogName AND PostID = @postID AND IFNULL(timestamp, 0) <> @timestamp";
|
||||
string sql = "UPDATE Notes SET TimeStamp = @timestamp, DateModified = @dateModified " +
|
||||
"WHERE RootBlogId = (SELECT BlogId FROM BlogNames WHERE BlogName = @rootBlogName) " +
|
||||
"AND NoteBlogId = (SELECT BlogId FROM BlogNames WHERE BlogName = @noteBlogName) " +
|
||||
"AND PostID = @postID AND IFNULL(TimeStamp, 0) <> @timestamp";
|
||||
using (SQLiteCommand command = new SQLiteCommand(sql, connection))
|
||||
{
|
||||
command.Parameters.AddWithValue("@timestamp", timestamp);
|
||||
@@ -1542,7 +1674,7 @@ namespace URLNotesGrabberCORE
|
||||
catch (Exception ex)
|
||||
{
|
||||
// Breakpoint here
|
||||
if (ex.Message != "constraint failed\r\nUNIQUE constraint failed: Notes.RootBlogName, Notes.PostID, Notes.TimeStamp, Notes.Type, Notes.NoteBlogName")
|
||||
if (!IsNotesDuplicateKey(ex))
|
||||
{
|
||||
Console.WriteLine(ex.Message);
|
||||
Console.WriteLine("^^^^^ - SHORTCUT");
|
||||
@@ -1552,9 +1684,10 @@ namespace URLNotesGrabberCORE
|
||||
return false;
|
||||
}
|
||||
|
||||
public static void UpdatePost(string blogName, long postID, string reblogURL, string postDate, string postURL, string slug, string reblogKey, string reblogName, string summary, string quote, string body, string tags, string link, string photoURL, string photoCaption, string downloadedFiles, string audioCaption, string question, string answer, string title, bool hasImage, bool byLikes = false, string? DBPath = null, string? rootBlogName = null, string? rootURL = null)
|
||||
public static void UpdatePost(string blogName, long postID, string reblogURL, string postDate, string postURL, string slug, string reblogKey, string reblogName, string summary, string quote, string body, string tags, string link, string photoURL, string photoCaption, string downloadedFiles, string audioCaption, string question, string answer, string title, bool hasImage, bool byLikes = false, string? DBPath = null, string? rootBlogName = null, string? rootURL = null, string? postType = null)
|
||||
{
|
||||
DBPath ??= GetDefaultDbPath();
|
||||
postType = PostTypes.Normalize(postType);
|
||||
|
||||
SQLiteConnection connection;
|
||||
bool ownsConnection;
|
||||
@@ -1602,6 +1735,10 @@ namespace URLNotesGrabberCORE
|
||||
sql += "RootBlogName = CASE WHEN @rootBlogName IS NULL OR @rootBlogName = '' OR @rootBlogName = '.' THEN RootBlogName ELSE @rootBlogName END, ";
|
||||
sql += "RootURL = CASE WHEN @rootURL IS NULL OR @rootURL = '' OR @rootURL = '.' THEN RootURL ELSE @rootURL END, ";
|
||||
sql += "hasImage = @hasImage, ";
|
||||
// Fill in a missing type, never overwrite one. A type derived by --ingest from a
|
||||
// real export filename is authoritative; this path's type is only as good as the
|
||||
// folder it was crawled from, so it must not win over an existing value.
|
||||
sql += "PostType = IFNULL(PostType, @postType), ";
|
||||
sql += "ByLikes = MAX(IFNULL(ByLikes, 0), @byLikes) ";
|
||||
sql += " WHERE BlogName = @BlogName AND PostID = @PostID AND (";
|
||||
sql += "(@postDate <> '.' AND IFNULL(postDate, '') <> @postDate) OR ";
|
||||
@@ -1625,7 +1762,10 @@ namespace URLNotesGrabberCORE
|
||||
sql += "IFNULL(hasImage, 0) <> @hasImage OR ";
|
||||
sql += "(@byLikes = 1 AND IFNULL(ByLikes, 0) = 0) OR ";
|
||||
sql += "((@rootBlogName IS NOT NULL AND @rootBlogName <> '' AND @rootBlogName <> '.') AND IFNULL(RootBlogName, '') <> @rootBlogName) OR ";
|
||||
sql += "((@rootURL IS NOT NULL AND @rootURL <> '' AND @rootURL <> '.') AND IFNULL(RootURL, '') <> @rootURL)";
|
||||
sql += "((@rootURL IS NOT NULL AND @rootURL <> '' AND @rootURL <> '.') AND IFNULL(RootURL, '') <> @rootURL) OR ";
|
||||
// Without this the SET above is unreachable for a row whose content is already
|
||||
// current: the UPDATE would not fire, and the type would stay NULL forever.
|
||||
sql += "(PostType IS NULL AND @postType IS NOT NULL)";
|
||||
sql += ")";
|
||||
|
||||
using (SQLiteCommand command = new SQLiteCommand(sql, connection))
|
||||
@@ -1653,6 +1793,7 @@ namespace URLNotesGrabberCORE
|
||||
command.Parameters.AddWithValue("@rootURL", string.IsNullOrWhiteSpace(rootURL) ? (object)DBNull.Value : rootURL);
|
||||
command.Parameters.AddWithValue("@hasImage", hasImage ? 1 : 0);
|
||||
command.Parameters.AddWithValue("@byLikes", byLikes ? 1 : 0);
|
||||
command.Parameters.AddWithValue("@postType", (object?)postType ?? DBNull.Value);
|
||||
command.Parameters.AddWithValue("@BlogName", blogName);
|
||||
command.Parameters.AddWithValue("@PostID", postID);
|
||||
|
||||
@@ -1828,10 +1969,15 @@ namespace URLNotesGrabberCORE
|
||||
{
|
||||
connection.Open();
|
||||
|
||||
//string sql = "UPDATE Notes SET replyText = @replyText WHERE rootBlogName = @rootBlogName AND PostID = @PostID AND noteBlogName = @noteBlogName AND TimeStamp = @TimeStamp AND Type = 'reply'";
|
||||
// Match on (noteBlogName, TimeStamp ±5s) only - a reply by a given blog at a given timestamp is the same reply across the original post and every reblog of it, so this fans out across reblog chains in one shot. Tolerance absorbs the ~1s drift between what -collect stored and what mode=conversation returns now.
|
||||
// Only fan out to rows that match the SELECT criteria in GetRepliesWithFilledText (NULL/empty/legacy-'.'). Never overwrite '?' (confirmed-empty) or already-fetched text.
|
||||
string sql = "UPDATE Notes SET replyText = @replyText, DateModified = @dateModified WHERE noteBlogName = @noteBlogName AND ABS(TimeStamp - @TimeStamp) <= 5 AND Type = 'reply' AND (replyText IS NULL OR replyText = '' OR replyText = '.') AND (replyText IS NULL OR replyText <> @replyText)";
|
||||
// The ABS() term cannot use an index on TimeStamp, before or after the integer schema; the NoteBlogId probe is what keeps this off a full scan.
|
||||
string sql = "UPDATE Notes SET replyText = @replyText, DateModified = @dateModified " +
|
||||
"WHERE NoteBlogId = (SELECT BlogId FROM BlogNames WHERE BlogName = @noteBlogName) " +
|
||||
"AND ABS(TimeStamp - @TimeStamp) <= 5 " +
|
||||
"AND TypeId = (SELECT TypeId FROM NoteTypes WHERE Type = 'reply') " +
|
||||
"AND (replyText IS NULL OR replyText = '' OR replyText = '.') " +
|
||||
"AND (replyText IS NULL OR replyText <> @replyText)";
|
||||
using (SQLiteCommand command = new SQLiteCommand(sql, connection))
|
||||
{
|
||||
command.Parameters.AddWithValue("@replyText", replyText ?? "?");
|
||||
@@ -1872,7 +2018,11 @@ namespace URLNotesGrabberCORE
|
||||
{
|
||||
connection.Open();
|
||||
|
||||
string sql = "UPDATE Notes SET replyText = @replyText, DateModified = @dateModified WHERE rootBlogName = @rootBlogName AND PostID = @PostID AND Type = 'reply' AND IFNULL(replyText, '.') <> @replyText";
|
||||
string sql = "UPDATE Notes SET replyText = @replyText, DateModified = @dateModified " +
|
||||
"WHERE RootBlogId = (SELECT BlogId FROM BlogNames WHERE BlogName = @rootBlogName) " +
|
||||
"AND PostID = @PostID " +
|
||||
"AND TypeId = (SELECT TypeId FROM NoteTypes WHERE Type = 'reply') " +
|
||||
"AND IFNULL(replyText, '.') <> @replyText";
|
||||
using (SQLiteCommand command = new SQLiteCommand(sql, connection))
|
||||
{
|
||||
command.Parameters.AddWithValue("@replyText", replyText ?? ".");
|
||||
@@ -1953,6 +2103,8 @@ namespace URLNotesGrabberCORE
|
||||
cmd.ExecuteNonQuery();
|
||||
Console.WriteLine("[Migration] Added PostType column to Posts table");
|
||||
}
|
||||
|
||||
BackfillMissingPostTypes(connection);
|
||||
}
|
||||
catch (Exception ex)
|
||||
{
|
||||
@@ -1960,6 +2112,65 @@ namespace URLNotesGrabberCORE
|
||||
}
|
||||
}
|
||||
|
||||
/// <summary>
|
||||
/// Types rows that carry no PostType, inferring it from which content columns they hold.
|
||||
///
|
||||
/// These are posts harvested from notes and likes rather than read out of a TumblThree
|
||||
/// export, so no filename ever described them and --ingest can never reach them: it only
|
||||
/// types a post it meets inside a real .txt. Content is the only signal they have.
|
||||
///
|
||||
/// Runs on every migration pass and is idempotent -- it only touches PostType IS NULL,
|
||||
/// so a row typed once is never revisited. Rows whose columns give no signal at all stay
|
||||
/// NULL and are skipped by OutputMode.
|
||||
///
|
||||
/// Mirrors PostTypes.InferFromContent; the two must agree. Notably HasImage is not
|
||||
/// consulted, because most text posts carry it.
|
||||
/// </summary>
|
||||
private static void BackfillMissingPostTypes(SQLiteConnection connection)
|
||||
{
|
||||
const string set = @"
|
||||
UPDATE Posts SET PostType = CASE
|
||||
WHEN Has(Question) AND Has(Answer) THEN 'answers'
|
||||
WHEN Has(Quote) THEN 'quotes'
|
||||
WHEN Has(Link) THEN 'links'
|
||||
WHEN Has(AudioCaption) THEN 'audios'
|
||||
WHEN Has(Body) THEN 'texts'
|
||||
WHEN Has(PhotoURL) OR Has(PhotoCaption) THEN 'images'
|
||||
ELSE NULL END
|
||||
WHERE PostType IS NULL";
|
||||
|
||||
// SQLite has no user-defined predicate here, so expand the "field supplied" test
|
||||
// ("." is the not-supplied sentinel used throughout the export format) inline.
|
||||
string sql = System.Text.RegularExpressions.Regex.Replace(
|
||||
set, @"Has\((\w+)\)", "TRIM(IFNULL($1, '')) NOT IN ('', '.')");
|
||||
|
||||
try
|
||||
{
|
||||
long before;
|
||||
using (var count = new SQLiteCommand("SELECT COUNT(*) FROM Posts WHERE PostType IS NULL", connection))
|
||||
before = Convert.ToInt64(count.ExecuteScalar());
|
||||
|
||||
if (before == 0) return;
|
||||
|
||||
int changed;
|
||||
using (var cmd = new SQLiteCommand(sql, connection))
|
||||
changed = cmd.ExecuteNonQuery();
|
||||
|
||||
long after;
|
||||
using (var count = new SQLiteCommand("SELECT COUNT(*) FROM Posts WHERE PostType IS NULL", connection))
|
||||
after = Convert.ToInt64(count.ExecuteScalar());
|
||||
|
||||
if (changed > 0 || after != before)
|
||||
Console.WriteLine($"[Migration] Backfilled PostType for {before - after} post(s); {after} still untyped (no content signal).");
|
||||
}
|
||||
catch (Exception ex)
|
||||
{
|
||||
// A failed backfill must not stop the run: untyped rows are skipped on export,
|
||||
// which is inconvenient, not corrupting.
|
||||
Console.WriteLine($"[Migration] PostType backfill failed: {ex.Message}");
|
||||
}
|
||||
}
|
||||
|
||||
// INSERT-or-UPDATE for a post arriving from a Tumblr text-file export.
|
||||
// On collision, only content columns + PostType + DateModified are updated;
|
||||
// engagement columns (ByLikes, RootBlogName, RootURL, HasNotesGathered, NotFound,
|
||||
@@ -1990,6 +2201,10 @@ namespace URLNotesGrabberCORE
|
||||
string? DBPath = null)
|
||||
{
|
||||
DBPath ??= GetDefaultDbPath();
|
||||
// Central guarantee: whatever a caller believes, only a canonical type reaches the
|
||||
// column. PostType is used as an output filename, so this is the invariant that keeps
|
||||
// a stray value from becoming a stray file.
|
||||
postType = PostTypes.Normalize(postType);
|
||||
try { AddBlog(blogName, false, DBPath); } catch { }
|
||||
|
||||
SQLiteConnection connection;
|
||||
|
||||
@@ -85,7 +85,25 @@ namespace URLNotesGrabberCORE
|
||||
{
|
||||
string rawBlogName = Path.GetFileName(Path.GetDirectoryName(file) ?? "unknown");
|
||||
string blogName = Regex.Replace(rawBlogName, @"_\d+$", "");
|
||||
string postType = Path.GetFileNameWithoutExtension(file);
|
||||
|
||||
// The filename becomes the row's PostType, and PostType later becomes an
|
||||
// output filename -- so an unrecognized name here would mint a new type and
|
||||
// a new file from any stray .txt that happens to sit in the tree. Only the
|
||||
// eight real export files are ingestable.
|
||||
//
|
||||
// This is also what breaks the Unknown.txt cycle: OutputMode used to write
|
||||
// untyped rows to Unknown.txt, and this scan would read it straight back
|
||||
// and stamp those rows with the literal type "Unknown", making the file
|
||||
// regenerate itself forever.
|
||||
string? resolvedPostType = PostTypes.FromFileName(file);
|
||||
if (resolvedPostType == null)
|
||||
{
|
||||
filesSkipped++;
|
||||
continue;
|
||||
}
|
||||
// Non-nullable from here so the local Flush() below stays warning-clean:
|
||||
// nullable flow analysis does not reach into local functions.
|
||||
string postType = resolvedPostType;
|
||||
|
||||
if (targetBlog != null && !string.Equals(blogName, targetBlog, StringComparison.OrdinalIgnoreCase))
|
||||
{
|
||||
|
||||
@@ -117,7 +117,12 @@ namespace URLNotesGrabberCORE
|
||||
question: reader.IsDBNull(18) ? null : reader.GetString(18),
|
||||
answer: reader.IsDBNull(19) ? null : reader.GetString(19),
|
||||
title: reader.IsDBNull(20) ? null : reader.GetString(20),
|
||||
postType: reader.IsDBNull(21) ? null : reader.GetString(21),
|
||||
// A legacy Posts.db predating the PostType column hands back NULL
|
||||
// here, and on the INSERT branch that NULL is stored -- reseeding
|
||||
// exactly the untyped rows the backfill exists to clear. Normalize
|
||||
// so an unrecognized legacy value cannot become a filename either;
|
||||
// the backfill types whatever comes through as null.
|
||||
postType: PostTypes.Normalize(reader.IsDBNull(21) ? null : reader.GetString(21)),
|
||||
hasImage: hasImage);
|
||||
postsUpserted++;
|
||||
if (postsUpserted % 500 == 0)
|
||||
|
||||
@@ -65,10 +65,20 @@ namespace URLNotesGrabberCORE
|
||||
var posts = DataAccess.GetAllPostsForBlog(blogName);
|
||||
Console.WriteLine($" Found {posts.Count} post(s) for this blog.");
|
||||
|
||||
var grouped = posts.GroupBy(p => p.PostType ?? "Unknown");
|
||||
// A post's type becomes a filename, so only a recognized type may be written. The
|
||||
// old `PostType ?? "Unknown"` invented Unknown.txt for untyped rows, which --ingest
|
||||
// then read back as a type named "Unknown" -- the two regenerated each other.
|
||||
// Untyped rows are skipped instead: after the backfill these are only rows with no
|
||||
// content signal at all, so nothing meaningful is lost, and nothing is invented.
|
||||
var typed = posts.Where(p => PostTypes.Normalize(p.PostType) != null).ToList();
|
||||
int untyped = posts.Count - typed.Count;
|
||||
if (untyped > 0)
|
||||
Console.WriteLine($" Skipping {untyped} post(s) with no recognized PostType.");
|
||||
|
||||
var grouped = typed.GroupBy(p => PostTypes.Normalize(p.PostType)!);
|
||||
foreach (var typeGroup in grouped)
|
||||
{
|
||||
string postType = typeGroup.Key ?? "Unknown";
|
||||
string postType = typeGroup.Key;
|
||||
string outputFilePath = Path.Combine(folder, $"{postType}.txt");
|
||||
var ordered = typeGroup.OrderBy(p => p.Date).ToList();
|
||||
Console.WriteLine($" Writing {ordered.Count} post(s) to {postType}.txt");
|
||||
|
||||
@@ -0,0 +1,123 @@
|
||||
using System;
|
||||
using System.Collections.Generic;
|
||||
using System.IO;
|
||||
|
||||
namespace URLNotesGrabberCORE
|
||||
{
|
||||
/// <summary>
|
||||
/// The single source of truth for Posts.PostType values.
|
||||
///
|
||||
/// PostType exists so --output can write one .txt per type. Because the type becomes a
|
||||
/// *filename*, an unvalidated value is not a cosmetic problem: it creates a file. That is
|
||||
/// how "Unknown.txt" came about -- OutputMode used `PostType ?? "Unknown"` as a filename,
|
||||
/// --ingest then read that file straight back and derived the literal type "Unknown" from
|
||||
/// its name, and the pair would have kept regenerating each other indefinitely.
|
||||
///
|
||||
/// So every path that produces a type routes through <see cref="Normalize"/>, which admits
|
||||
/// only the eight known names and returns null for anything else. A null type is safe:
|
||||
/// OutputMode skips those rows rather than inventing a file for them.
|
||||
/// </summary>
|
||||
public static class PostTypes
|
||||
{
|
||||
// The canonical set. These are exactly the TumblThree .txt basenames, which is what
|
||||
// makes an ingested filename usable as a type without translation.
|
||||
public const string Texts = "texts";
|
||||
public const string Answers = "answers";
|
||||
public const string Quotes = "quotes";
|
||||
public const string Links = "links";
|
||||
public const string Conversations = "conversations";
|
||||
public const string Images = "images";
|
||||
public const string Videos = "videos";
|
||||
public const string Audios = "audios";
|
||||
|
||||
private static readonly HashSet<string> Known = new HashSet<string>(
|
||||
new[] { Texts, Answers, Quotes, Links, Conversations, Images, Videos, Audios },
|
||||
StringComparer.OrdinalIgnoreCase);
|
||||
|
||||
// Tumblr's legacy post format (npf=false) names types in the singular. The likes API is
|
||||
// the one source that reports a type directly rather than via a filename, so it is the
|
||||
// only place this mapping is needed.
|
||||
private static readonly Dictionary<string, string> ApiTypeMap = new Dictionary<string, string>(StringComparer.OrdinalIgnoreCase)
|
||||
{
|
||||
["text"] = Texts,
|
||||
["photo"] = Images,
|
||||
["quote"] = Quotes,
|
||||
["link"] = Links,
|
||||
["chat"] = Conversations,
|
||||
["answer"] = Answers,
|
||||
["audio"] = Audios,
|
||||
["video"] = Videos,
|
||||
};
|
||||
|
||||
/// <summary>
|
||||
/// Returns the canonical type name, or null if the value is not one of the eight.
|
||||
/// Returning null rather than passing the value through is the whole point: an
|
||||
/// unrecognized string must never reach a filename.
|
||||
/// </summary>
|
||||
public static string? Normalize(string? candidate)
|
||||
{
|
||||
if (string.IsNullOrWhiteSpace(candidate)) return null;
|
||||
string trimmed = candidate.Trim();
|
||||
return Known.TryGetValue(trimmed, out string? canonical) ? canonical : null;
|
||||
}
|
||||
|
||||
/// <summary>
|
||||
/// Type for a post read out of a TumblThree export file, taken from the filename
|
||||
/// ("texts.txt" -> "texts"). Anything else in the folder -- README.txt, a stray
|
||||
/// triage file, or a previously written Unknown.txt -- normalizes to null and is
|
||||
/// rejected by the caller.
|
||||
/// </summary>
|
||||
public static string? FromFileName(string? path)
|
||||
{
|
||||
if (string.IsNullOrWhiteSpace(path)) return null;
|
||||
return Normalize(Path.GetFileNameWithoutExtension(path));
|
||||
}
|
||||
|
||||
/// <summary>
|
||||
/// Type for a post from the likes API, whose legacy-format `type` field is singular.
|
||||
/// Null when the field is absent or unrecognized -- the access is dynamic, so a missing
|
||||
/// field yields null at runtime rather than failing to compile.
|
||||
/// </summary>
|
||||
public static string? FromApiType(string? apiType)
|
||||
{
|
||||
if (string.IsNullOrWhiteSpace(apiType)) return null;
|
||||
return ApiTypeMap.TryGetValue(apiType.Trim(), out string? mapped) ? mapped : null;
|
||||
}
|
||||
|
||||
/// <summary>
|
||||
/// Last-resort type inferred from which content columns a row actually carries. Used
|
||||
/// only to backfill rows written before any type was recorded; a filename or an API
|
||||
/// type is always preferred over this.
|
||||
///
|
||||
/// The order matters and is derived from the already-typed rows, where the column
|
||||
/// signatures are effectively disjoint: answers carry Question+Answer and no Body,
|
||||
/// images carry photo columns and no Body, texts carry Body and no photo columns.
|
||||
///
|
||||
/// HasImage is deliberately NOT consulted: it is set on 12,420 of 19,828 known text
|
||||
/// posts, so it says nothing about the post's type.
|
||||
///
|
||||
/// conversations cannot be separated from texts this way -- both carry only Body -- so
|
||||
/// a chat post with no other signal is labelled texts. A later --ingest that meets the
|
||||
/// post in a real conversations.txt corrects it.
|
||||
/// </summary>
|
||||
public static string? InferFromContent(string? question, string? answer, string? quote,
|
||||
string? link, string? audioCaption, string? body, string? photoUrl, string? photoCaption)
|
||||
{
|
||||
if (HasValue(question) && HasValue(answer)) return Answers;
|
||||
if (HasValue(quote)) return Quotes;
|
||||
if (HasValue(link)) return Links;
|
||||
if (HasValue(audioCaption)) return Audios;
|
||||
if (HasValue(body)) return Texts;
|
||||
if (HasValue(photoUrl) || HasValue(photoCaption)) return Images;
|
||||
return null;
|
||||
}
|
||||
|
||||
// "." is the codebase-wide "field not supplied" sentinel in export records, so it
|
||||
// counts as absent here just as it does in UpdatePost's CASE guards.
|
||||
private static bool HasValue(string? value)
|
||||
{
|
||||
if (string.IsNullOrWhiteSpace(value)) return false;
|
||||
return value.Trim() != ".";
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -148,6 +148,12 @@ namespace URLNotesGrabberCORE
|
||||
if (args.Length == 0) //Traverse folder structure to add posts and thus blogs to DB
|
||||
{
|
||||
int postsAdded = 0;
|
||||
// This is the mode that actually gets run day to day, so the schema migration and
|
||||
// the PostType backfill have to happen here too. They used to hang off --ingest,
|
||||
// --output and friends only, which meant the untyped rows this traversal creates
|
||||
// could sit unrepaired indefinitely while the one command everyone runs skipped
|
||||
// the fix entirely. Idempotent, so paying it on every run costs nothing.
|
||||
DataAccess.EnsureTTFileHelperColumnsExist();
|
||||
try
|
||||
{
|
||||
DataAccess.EnableImportModePragmas();
|
||||
@@ -220,7 +226,7 @@ namespace URLNotesGrabberCORE
|
||||
|
||||
if (args.Length < 2)
|
||||
{
|
||||
Console.WriteLine("--Expected WITHOUTNOTESONLY (0, 1) [OPTIONAL: BEFOREDATE]--");
|
||||
Console.WriteLine("--Expected WITHOUTNOTESONLY (0, 1) [OPTIONAL: BEFOREDATE] [OPTIONAL: BLOGNAME]--");
|
||||
exitCode = 2;
|
||||
break;
|
||||
}
|
||||
@@ -242,28 +248,63 @@ namespace URLNotesGrabberCORE
|
||||
Console.WriteLine("Without Notes Only: {0}\t{1}", withoutNotesOnly, args[1]);
|
||||
}
|
||||
|
||||
// Parse optional beforeDate parameter
|
||||
if (args.Length >= 3 && !string.IsNullOrEmpty(args[2]))
|
||||
// Trailing arguments are the optional cutoff date and the optional blog filter, in
|
||||
// either order. A token that parses as a date is the cutoff; anything else is a blog
|
||||
// name -- which is why an unparseable token is no longer an error here. "--blog=name"
|
||||
// forces the blog reading for the rare name that would otherwise parse as a date.
|
||||
string? collectBlogName = null;
|
||||
bool badCollectArg = false;
|
||||
|
||||
for (int i = 2; i < args.Length; i++)
|
||||
{
|
||||
if (DateTime.TryParse(args[2], out DateTime parsedDate))
|
||||
string arg = args[i];
|
||||
if (string.IsNullOrWhiteSpace(arg))
|
||||
continue;
|
||||
|
||||
if (arg.StartsWith("--blog=", StringComparison.OrdinalIgnoreCase))
|
||||
{
|
||||
collectBlogName = arg.Substring("--blog=".Length);
|
||||
if (string.IsNullOrWhiteSpace(collectBlogName))
|
||||
{
|
||||
Console.WriteLine("ERROR: --blog= requires a blog name");
|
||||
badCollectArg = true;
|
||||
break;
|
||||
}
|
||||
}
|
||||
else if (!explicitDateSupplied && DateTime.TryParse(arg, out DateTime parsedDate))
|
||||
{
|
||||
beforeDate = parsedDate;
|
||||
explicitDateSupplied = true;
|
||||
Console.WriteLine($"Filter: Collecting notes for posts with NotesGatheredDateTime < {beforeDate}");
|
||||
}
|
||||
else if (collectBlogName == null)
|
||||
{
|
||||
collectBlogName = arg;
|
||||
}
|
||||
else
|
||||
{
|
||||
Console.WriteLine($"ERROR: Invalid date format '{args[2]}'");
|
||||
exitCode = 2;
|
||||
Console.WriteLine($"ERROR: Unexpected argument '{arg}'");
|
||||
badCollectArg = true;
|
||||
break;
|
||||
}
|
||||
}
|
||||
|
||||
if (badCollectArg)
|
||||
{
|
||||
exitCode = 2;
|
||||
break;
|
||||
}
|
||||
|
||||
if (collectBlogName != null)
|
||||
Console.WriteLine($"Filter: Collecting notes for posts by '{collectBlogName}' only");
|
||||
|
||||
// Mode 0 (full re-check) with no explicit date is a *managed* run: freeze the cutoff and
|
||||
// persist it so an interrupted run resumes against the same cutoff and a completed run stops
|
||||
// instead of restarting. Mode 1 and explicit-date runs keep their existing behavior.
|
||||
// instead of restarting. Mode 1, explicit-date and blog-scoped runs keep their existing
|
||||
// behavior -- a single blog covers a slice of the worklist, so letting it write the shared
|
||||
// run state would mark the whole re-check complete after collecting one blog.
|
||||
bool managedCollectRun = false;
|
||||
if (!withoutNotesOnly && !explicitDateSupplied)
|
||||
if (!withoutNotesOnly && !explicitDateSupplied && collectBlogName == null)
|
||||
{
|
||||
DataAccess.EnsureCollectRunStateTableExists();
|
||||
var runState = DataAccess.GetCollectRunState();
|
||||
@@ -281,7 +322,7 @@ namespace URLNotesGrabberCORE
|
||||
managedCollectRun = true;
|
||||
}
|
||||
|
||||
exitCode = CollectNotes(settings.GetValue<string>("PathOutput"), withoutNotesOnly, beforeDate, managedCollectRun).GetAwaiter().GetResult();
|
||||
exitCode = CollectNotes(settings.GetValue<string>("PathOutput"), withoutNotesOnly, beforeDate, managedCollectRun, collectBlogName).GetAwaiter().GetResult();
|
||||
break;
|
||||
|
||||
case "--blogsR": //collect notes from all posts
|
||||
@@ -413,7 +454,7 @@ namespace URLNotesGrabberCORE
|
||||
|
||||
Console.WriteLine("--blogs\t For each Blog in DB, write blogname to file");
|
||||
|
||||
Console.WriteLine("--collect [0|1] [datetime]\t Collect Notes from API. 1=only posts without notes. 0=full re-check of all posts: a single resumable pass (interrupt & relaunch to resume; stops when complete, retrigger for a new pass). Optional datetime overrides the cutoff and runs as a one-off (bypasses resume tracking).");
|
||||
Console.WriteLine("--collect [0|1] [datetime] [blogname]\t Collect Notes from API. 1=only posts without notes. 0=full re-check of all posts: a single resumable pass (interrupt & relaunch to resume; stops when complete, retrigger for a new pass). Optional datetime overrides the cutoff and runs as a one-off (bypasses resume tracking). Optional blogname restricts the run to that blog (exact, case-sensitive match) and also runs as a one-off; e.g. \"--collect 1 zomb-eh\". datetime and blogname may be given in either order - use --blog=name if a blog name would otherwise parse as a date.");
|
||||
|
||||
Console.WriteLine("--blogsR\t For each Note that is a REPLY, write blogname to file ");
|
||||
|
||||
@@ -1012,6 +1053,15 @@ namespace URLNotesGrabberCORE
|
||||
string reblogKey = post.reblog_key?.ToString() ?? ".";
|
||||
string link = ".";
|
||||
|
||||
// No file backs a liked post, so the filename trick used everywhere
|
||||
// else cannot apply here. GrabLikes requests npf=false, and in the
|
||||
// legacy format `type` is the discriminator that decides which content
|
||||
// fields a post carries -- singular there, mapped to our plural names.
|
||||
// liked_posts is List<dynamic>, so this is resolved at runtime and a
|
||||
// missing field yields null rather than a compile error; an absent or
|
||||
// unrecognized value leaves the type NULL instead of guessing.
|
||||
string? apiPostType = PostTypes.FromApiType(post.type?.ToString() as string);
|
||||
|
||||
// Only insert if any of the data contains strings from ContainsList
|
||||
bool shouldInsert = false;
|
||||
string matchedFieldName = string.Empty;
|
||||
@@ -1055,7 +1105,8 @@ if (shouldInsert)
|
||||
DataAccess.AddPost(authorBlog, postID, reblogURL, date, postURL, slug, reblogKey,
|
||||
reblogName, summary, quote, body, tags, link, photoURL,
|
||||
photoCaption, downloadedFiles, audioCaption, question, answer,
|
||||
title, hasImage, true, rootBlogName: rootBlogName, rootURL: rootURL);
|
||||
title, hasImage, true, rootBlogName: rootBlogName, rootURL: rootURL,
|
||||
postType: apiPostType);
|
||||
}
|
||||
}
|
||||
|
||||
@@ -1294,9 +1345,17 @@ if (shouldInsert)
|
||||
// blipping on one post. Past this, skipping post-by-post would just hammer a closed door.
|
||||
const int MaxConsecutiveTransient = 10;
|
||||
|
||||
static async Task<int> CollectNotes(string outPath, bool withoutNotesOnly = true, DateTime? beforeDate = null, bool managedRun = false)
|
||||
static async Task<int> CollectNotes(string outPath, bool withoutNotesOnly = true, DateTime? beforeDate = null, bool managedRun = false, string? blogName = null)
|
||||
{
|
||||
List<Tuple<string, long, long, long>> posts = DataAccess.GetPosts(withoutNotesOnly, beforeDate);
|
||||
List<Tuple<string, long, long, long>> posts = DataAccess.GetPosts(withoutNotesOnly, beforeDate, blogName);
|
||||
|
||||
if (posts.Count == 0 && !string.IsNullOrWhiteSpace(blogName))
|
||||
{
|
||||
// BlogName is matched exactly, so a typo or a case mismatch looks identical to "nothing
|
||||
// left to collect". Say so rather than reporting a silent, instant success.
|
||||
Console.WriteLine($"No posts to collect for blog '{blogName}'. Either it is fully collected, or the name does not match a stored blog (the match is case-sensitive).");
|
||||
return 0;
|
||||
}
|
||||
|
||||
// Posts attempted (with a definitive, non-throttle result) during *this* process. Guarantees a single
|
||||
// attempt pass: once every remaining post has been attempted, the loop stops instead of spinning on a
|
||||
@@ -1392,7 +1451,7 @@ if (shouldInsert)
|
||||
}
|
||||
|
||||
// Re-fetch the updated list after processing the current post
|
||||
posts = DataAccess.GetPosts(withoutNotesOnly, beforeDate);
|
||||
posts = DataAccess.GetPosts(withoutNotesOnly, beforeDate, blogName);
|
||||
}
|
||||
}
|
||||
|
||||
@@ -1465,7 +1524,15 @@ if (shouldInsert)
|
||||
string normalizedDirectoryName = NormalizeBlogFolderName(new DirectoryInfo(path).Name);
|
||||
bool isAtOrAfterStart = string.IsNullOrWhiteSpace(startFromBlogName) || string.Compare(normalizedDirectoryName, startFromBlogName, StringComparison.OrdinalIgnoreCase) >= 0;
|
||||
|
||||
// The filename is the post type ("texts.txt" -> "texts"), so only the eight
|
||||
// known export files are post sources. Everything else in a blog folder is
|
||||
// either not a post file at all (README.txt, url lists, triage scratch) or
|
||||
// is our own derived output -- Unknown.txt above all, which must never be
|
||||
// read back in as a source or it perpetuates itself.
|
||||
string? filePostType = PostTypes.FromFileName(file);
|
||||
|
||||
if (file.EndsWith(".txt", StringComparison.OrdinalIgnoreCase)
|
||||
&& filePostType != null
|
||||
&& (string.IsNullOrEmpty(blogName) || path.IndexOf(blogName, StringComparison.OrdinalIgnoreCase) >= 0)
|
||||
&& isAtOrAfterStart)
|
||||
{
|
||||
@@ -1498,7 +1565,7 @@ if (shouldInsert)
|
||||
DataAccess.AddPost(curDir, long.Parse(reblog.postID), reblog.reblogURL, reblog.date, reblog.postURL, reblog.slug, reblog.reblogKey,
|
||||
reblog.reblogName, reblog.summary, reblog.quote, reblog.body, reblog.tags, reblog.link, reblog.photoURL,
|
||||
reblog.photoCaption, reblog.downloadedFiles, reblog.audioCaption, reblog.question, reblog.answer,
|
||||
reblog.title, false, rootURL: reblog.rootURL);
|
||||
reblog.title, false, rootURL: reblog.rootURL, postType: filePostType);
|
||||
recordImportStopwatch.Stop();
|
||||
|
||||
postsAdded++;
|
||||
@@ -1646,7 +1713,7 @@ if (shouldInsert)
|
||||
DataAccess.AddPost(curDir, long.Parse(reblog.postID), reblog.reblogURL, reblog.date, reblog.postURL, reblog.slug, reblog.reblogKey,
|
||||
reblog.reblogName, reblog.summary, reblog.quote, reblog.body, reblog.tags, reblog.link, reblog.photoURL,
|
||||
reblog.photoCaption, reblog.downloadedFiles, reblog.audioCaption, reblog.question, reblog.answer,
|
||||
reblog.title, true, rootURL: reblog.rootURL);
|
||||
reblog.title, true, rootURL: reblog.rootURL, postType: filePostType);
|
||||
recordImportStopwatch.Stop();
|
||||
|
||||
postsAdded++;
|
||||
|
||||
@@ -313,9 +313,17 @@ code to this table's contents, so prefer the join in anything long-lived.
|
||||
|
||||
## Porting to the integer schema
|
||||
|
||||
Everything here was checked against the live 148 MB file. There are roughly 14 affected
|
||||
call sites in `DataAccess.cs` and 16 in `RolodexRepository.cs`. TumblThree needs no
|
||||
changes — its single statement touches `Blogs.IsActive` and `BlogName` only.
|
||||
Everything here was checked against the live 148 MB file. There were 14 affected call
|
||||
sites in `DataAccess.cs` and 16 in `RolodexRepository.cs`. TumblThree needs no changes —
|
||||
its single statement touches `Blogs.IsActive` and `BlogName` only.
|
||||
|
||||
**`DataAccess.cs` is ported.** All 14 sites now read the integer schema, `AddNote`
|
||||
registers names and types before inserting, and `verify-db-schema.sql` reports a
|
||||
pre-migration file rather than letting the app fail on it. `RolodexRepository.cs` lives in
|
||||
the [Rolodex](https://git.basso.land/jim/Rolodex) repository and is not covered by that
|
||||
work. One site was dropped rather than translated: the `LEFT JOIN Notes` in `GetPosts`
|
||||
selected nothing and was collapsed by the query's own `GROUP BY`, so it could not affect
|
||||
the result.
|
||||
|
||||
### Column mapping
|
||||
|
||||
@@ -426,7 +434,11 @@ UNIQUE constraint failed: Notes.RootBlogName, Notes.PostID, Notes.TimeStamp, Not
|
||||
|
||||
at two call sites to decide whether to swallow an exception. SQLite now emits the *new*
|
||||
column names, so those comparisons no longer match and real errors will surface where
|
||||
they used to be silently ignored — or vice versa. Both sites need updating.
|
||||
they used to be silently ignored — or vice versa.
|
||||
|
||||
Both sites now go through `IsNotesDuplicateKey` in `DataAccess.cs`, which matches on
|
||||
`UNIQUE constraint failed` plus `Notes.` rather than on the column list. A literal
|
||||
comparison is what broke here; the next rename should not break it again.
|
||||
|
||||
### Updating notes
|
||||
|
||||
@@ -526,7 +538,7 @@ Crawler bookkeeping. Rolodex ignores all of these.
|
||||
`DataAccess.cs` joins on it to decide what to collect:
|
||||
|
||||
```sql
|
||||
-- as it will read after the integer-schema port; see the porting guide above
|
||||
-- shape only; the ported GetBlogs joins Blogs directly on BlogId and needs no BlogNames hop
|
||||
SELECT bn.BlogName, count(*)
|
||||
FROM Notes n
|
||||
JOIN Blogs b ON b.BlogId = n.NoteBlogId
|
||||
|
||||
+54
-9
@@ -79,18 +79,35 @@ WITH expected(tbl, col, alter_stmt) AS (
|
||||
('Blogs','LikesLastRefreshed', 'ALTER TABLE Blogs ADD COLUMN LikesLastRefreshed INTEGER DEFAULT 0;'),
|
||||
('Blogs','LikesLastNewCount', 'ALTER TABLE Blogs ADD COLUMN LikesLastNewCount INTEGER DEFAULT 0;'),
|
||||
('Blogs','TTFolderPath', 'ALTER TABLE Blogs ADD COLUMN TTFolderPath TEXT;'),
|
||||
-- Blogs.BlogId (2026-08-07) is the single-hop join key into Notes. Deliberately NOT
|
||||
-- auto-fixable: an added-but-empty BlogId makes every engagement join return zero
|
||||
-- rows silently, which is worse than the hard error a missing column gives.
|
||||
('Blogs','BlogId', 'MANUAL REVIEW - see query 1d: run normalize-notes.sql'),
|
||||
|
||||
-- Notes (base columns: manual review if missing)
|
||||
('Notes','RootBlogName', 'MANUAL REVIEW - base/PK column missing'),
|
||||
-- Integer IDs since 2026-08-07. RootBlogName/NoteBlogName/Type are GONE, not renamed
|
||||
-- in place -- a backup that still has them needs normalize-notes.sql, not an ALTER.
|
||||
-- Query 1d below reports exactly that case.
|
||||
('Notes','RootBlogId', 'MANUAL REVIEW - see query 1d: pre-2026-08-07 name schema, or damaged'),
|
||||
('Notes','PostID', 'MANUAL REVIEW - base/PK column missing'),
|
||||
('Notes','NoteBlogName', 'MANUAL REVIEW - base/PK column missing'),
|
||||
('Notes','NoteBlogId', 'MANUAL REVIEW - see query 1d: pre-2026-08-07 name schema, or damaged'),
|
||||
('Notes','TimeStamp', 'MANUAL REVIEW - base/PK column missing'),
|
||||
('Notes','Type', 'MANUAL REVIEW - base/PK column missing'),
|
||||
('Notes','TypeId', 'MANUAL REVIEW - see query 1d: pre-2026-08-07 name schema, or damaged'),
|
||||
('Notes','DatetimeCrawled', 'MANUAL REVIEW - base column missing'),
|
||||
('Notes','DateModified', 'MANUAL REVIEW - base column missing'),
|
||||
('Notes','DateCreated', 'MANUAL REVIEW - base column missing'),
|
||||
-- Notes (additive migration column, auto-fixable)
|
||||
('Notes','replyText', 'ALTER TABLE Notes ADD COLUMN replyText TEXT DEFAULT ''.'';'),
|
||||
-- No DEFAULT: the migrated schema dropped it, so new rows get NULL rather than a
|
||||
-- placeholder. EnsureReplyTextColumnExists in DataAccess.cs adds it the same way.
|
||||
('Notes','replyText', 'ALTER TABLE Notes ADD COLUMN replyText TEXT;'),
|
||||
|
||||
-- BlogNames / NoteTypes (the lookup tables Notes resolves its IDs through, 2026-08-07).
|
||||
-- Not auto-fixable: an empty BlogNames does not mean "add the table", it means the
|
||||
-- Notes rows have nothing to resolve against. Rebuild with normalize-notes.sql.
|
||||
('BlogNames','BlogId', 'MANUAL REVIEW - see query 1d: run normalize-notes.sql'),
|
||||
('BlogNames','BlogName', 'MANUAL REVIEW - see query 1d: run normalize-notes.sql'),
|
||||
('NoteTypes','TypeId', 'MANUAL REVIEW - see query 1d: run normalize-notes.sql'),
|
||||
('NoteTypes','Type', 'MANUAL REVIEW - see query 1d: run normalize-notes.sql'),
|
||||
|
||||
-- DailyAPICount (base columns)
|
||||
('DailyAPICount','Date', 'MANUAL REVIEW - base/PK column missing'),
|
||||
@@ -106,6 +123,8 @@ actual(tbl, col) AS (
|
||||
SELECT 'Posts', name FROM pragma_table_info('Posts')
|
||||
UNION ALL SELECT 'Blogs', name FROM pragma_table_info('Blogs')
|
||||
UNION ALL SELECT 'Notes', name FROM pragma_table_info('Notes')
|
||||
UNION ALL SELECT 'BlogNames', name FROM pragma_table_info('BlogNames')
|
||||
UNION ALL SELECT 'NoteTypes', name FROM pragma_table_info('NoteTypes')
|
||||
UNION ALL SELECT 'DailyAPICount', name FROM pragma_table_info('DailyAPICount')
|
||||
UNION ALL SELECT 'ApiKeyPoolState', name FROM pragma_table_info('ApiKeyPoolState')
|
||||
UNION ALL SELECT 'ApiKeyPoolMeta', name FROM pragma_table_info('ApiKeyPoolMeta')
|
||||
@@ -127,7 +146,7 @@ ORDER BY (e.alter_stmt LIKE 'ALTER%') DESC, e.tbl, e.col;
|
||||
-- 1b. MISSING TABLES: expected tables that don't exist at all in this DB.
|
||||
-- Zero rows = good.
|
||||
WITH expected_tables(tbl) AS (
|
||||
VALUES ('Posts'),('Blogs'),('Notes'),('DailyAPICount'),
|
||||
VALUES ('Posts'),('Blogs'),('Notes'),('BlogNames'),('NoteTypes'),('DailyAPICount'),
|
||||
('ApiKeyPoolState'),('ApiKeyPoolMeta')
|
||||
)
|
||||
SELECT et.tbl AS missing_table
|
||||
@@ -157,10 +176,12 @@ WITH expected(tbl, col) AS (
|
||||
('Blogs','BlogName'),('Blogs','HasBeenOutput'),('Blogs','IsActive'),('Blogs','DateAdded'),
|
||||
('Blogs','ByLikes'),('Blogs','DateModified'),('Blogs','DateCreated'),('Blogs','LikesPulled'),
|
||||
('Blogs','LikesCursor'),('Blogs','LikesNewestTimestamp'),('Blogs','LikesLastRefreshed'),
|
||||
('Blogs','LikesLastNewCount'),('Blogs','TTFolderPath'),
|
||||
('Notes','RootBlogName'),('Notes','PostID'),('Notes','NoteBlogName'),('Notes','TimeStamp'),
|
||||
('Notes','Type'),('Notes','DatetimeCrawled'),('Notes','DateModified'),('Notes','DateCreated'),
|
||||
('Blogs','LikesLastNewCount'),('Blogs','TTFolderPath'),('Blogs','BlogId'),
|
||||
('Notes','RootBlogId'),('Notes','PostID'),('Notes','NoteBlogId'),('Notes','TimeStamp'),
|
||||
('Notes','TypeId'),('Notes','DatetimeCrawled'),('Notes','DateModified'),('Notes','DateCreated'),
|
||||
('Notes','replyText'),('Notes','IsActive'),
|
||||
('BlogNames','BlogId'),('BlogNames','BlogName'),
|
||||
('NoteTypes','TypeId'),('NoteTypes','Type'),
|
||||
('DailyAPICount','Date'),('DailyAPICount','APICount'),
|
||||
('ApiKeyPoolState','KeyName'),('ApiKeyPoolState','RetryUntil'),
|
||||
('ApiKeyPoolMeta','Id'),('ApiKeyPoolMeta','LastIndex')
|
||||
@@ -169,6 +190,8 @@ actual(tbl, col) AS (
|
||||
SELECT 'Posts', name FROM pragma_table_info('Posts')
|
||||
UNION ALL SELECT 'Blogs', name FROM pragma_table_info('Blogs')
|
||||
UNION ALL SELECT 'Notes', name FROM pragma_table_info('Notes')
|
||||
UNION ALL SELECT 'BlogNames', name FROM pragma_table_info('BlogNames')
|
||||
UNION ALL SELECT 'NoteTypes', name FROM pragma_table_info('NoteTypes')
|
||||
UNION ALL SELECT 'DailyAPICount', name FROM pragma_table_info('DailyAPICount')
|
||||
UNION ALL SELECT 'ApiKeyPoolState', name FROM pragma_table_info('ApiKeyPoolState')
|
||||
UNION ALL SELECT 'ApiKeyPoolMeta', name FROM pragma_table_info('ApiKeyPoolMeta')
|
||||
@@ -181,6 +204,25 @@ WHERE e.col IS NULL
|
||||
ORDER BY a.tbl, a.col;
|
||||
|
||||
|
||||
-- 1d. PRE-MIGRATION DATABASE: a backup from before 2026-08-07, when Notes still
|
||||
-- stored names. Zero rows = good.
|
||||
--
|
||||
-- This is the one failure SECTION 2 cannot fix. Notes.RootBlogName /
|
||||
-- NoteBlogName / Type were replaced by RootBlogId / NoteBlogId / TypeId
|
||||
-- resolving through BlogNames and NoteTypes -- a data migration, not an
|
||||
-- ADD COLUMN. There is no compatibility view, so the current code fails
|
||||
-- outright ("no such column: RootBlogId") against such a file.
|
||||
--
|
||||
-- Fix: run normalize-notes.sql against a COPY of the backup, then re-run
|
||||
-- SECTION 1. Do not hand-add the ID columns: they would be empty, and an
|
||||
-- empty NoteBlogId is indistinguishable from a note by blog #0.
|
||||
SELECT 'Notes still stores names -- run normalize-notes.sql on a copy' AS pre_migration_schema,
|
||||
group_concat(name, ', ') AS legacy_columns_found
|
||||
FROM pragma_table_info('Notes')
|
||||
WHERE lower(name) IN ('rootblogname','noteblogname','type')
|
||||
HAVING COUNT(*) > 0;
|
||||
|
||||
|
||||
-- ============================================================================
|
||||
-- SECTION 2 -- FIX (opt-in, additive only)
|
||||
--
|
||||
@@ -190,6 +232,9 @@ ORDER BY a.tbl, a.col;
|
||||
-- "duplicate column name" error and changes nothing -- just run the flagged
|
||||
-- subset. These are the 8 additive migration columns and nothing else; the
|
||||
-- likes high-water-mark reset is intentionally NOT included.
|
||||
--
|
||||
-- Nothing here addresses query 1d. The Notes integer schema is a data migration
|
||||
-- (normalize-notes.sql) and cannot be reached by adding columns.
|
||||
-- ============================================================================
|
||||
|
||||
-- ALTER TABLE Posts ADD COLUMN PostType TEXT;
|
||||
@@ -199,4 +244,4 @@ ORDER BY a.tbl, a.col;
|
||||
-- ALTER TABLE Blogs ADD COLUMN LikesLastRefreshed INTEGER DEFAULT 0;
|
||||
-- ALTER TABLE Blogs ADD COLUMN LikesLastNewCount INTEGER DEFAULT 0;
|
||||
-- ALTER TABLE Blogs ADD COLUMN TTFolderPath TEXT;
|
||||
-- ALTER TABLE Notes ADD COLUMN replyText TEXT DEFAULT '.';
|
||||
-- ALTER TABLE Notes ADD COLUMN replyText TEXT;
|
||||
|
||||
Reference in New Issue
Block a user