Author SHA1 Message Date
jimandClaude Opus 5 f41957fd2f feat(collect): let --force ignore the re-collect cooldown
--collect 1 <blog> returned an empty worklist whenever the periodic
re-queue branch was still inside its 3-day window, with no way to ask
for the re-collect early. GetPosts now composes that age predicate
conditionally, and the existing global --force flag - already "ignore
cooldown" for --likes - drives it.

Only the age gate drops: NotFound = 0, the IsActive filter and the blog
scoping still apply. The flag is a no-op for mode 0, which re-collects
every post regardless, and says so rather than pretending to act.

Co-Authored-By: Claude Opus 5 <[email protected]>
2026-08-23 04:34:45 -05:00
jimandClaude Sonnet 5 a28c5cc9ec fix(dbbrowser): rewrite saved queries for the Notes integer schema
Notes.RootBlogName/NoteBlogName/Type were replaced by RootBlogId/NoteBlogId/
TypeId (resolved via BlogNames and NoteTypes) back on 2026-08-07. The saved
DB Browser for SQLite queries in RERUN.sqbpro and TL.sqbpro still referenced
the old text columns and failed against the migrated TL.db.

Rewrote the 5 affected queries per TL.db.md's porting guide: Notes<->Blogs
joins go through Blogs.BlogId in one hop, Notes<->Posts joins route through
BlogNames (Posts has no BlogId), and type filters resolve through NoteTypes.
Verified read-only against the live TL.db -- all 5 execute without error.

Co-Authored-By: Claude Sonnet 5 <[email protected]>
2026-08-20 09:07:08 -05:00
jim d6a96f7885 Merge branch 'claude/posttype-noargs-backfill' into master 2026-08-20 08:15:54 -05:00
jimandClaude Opus 5 864b468d96 fix(posttype): run the migration and backfill in no-args mode
The backfill hung off --ingest, --output, --correct, --importposts and
--updatepaths, because those were where EnsureTTFileHelperColumnsExist was
already being called. But the no-argument traversal is the mode that
actually gets run day to day, and it called none of them -- so the command
used most often was the one command that never repaired an untyped row.

TraverseDirectory already types the posts it inserts, from the filename it
is reading. This closes the other half: the rows already sitting untyped
now get fixed by an ordinary run, with no separate maintenance command.

Called before BeginImportSession so it uses its own connection rather than
contending with the import session's, and before the traversal so existing
rows are typed first and newly inserted ones arrive already typed.

Co-Authored-By: Claude Opus 5 <[email protected]>
2026-08-20 08:15:54 -05:00
jim e6a3efba5b Merge branch 'claude/posttype-unknown-fix' into master 2026-08-20 08:01:01 -05:00
jimandClaude Opus 5 34da632e6a fix(posttype): stop Unknown.txt and type posts at their source
PostType becomes an output filename, so an unset or unvalidated value does
not stay a data problem -- it creates a file. OutputMode wrote untyped rows
to `PostType ?? "Unknown"`, IngestMode read that file back and derived the
literal type "Unknown" from its name, and the two would have regenerated
each other indefinitely.

Nothing was setting the type in the first place. AddPost -- the path every
notes/likes harvest goes through -- omitted PostType from its INSERT column
list entirely, so 1790 rows across 469 blogs had none. Ingest could never
repair them: it types a post only when it meets it inside a real export
file, and these posts appear in none.

Type at the source, from what each path actually knows:

- TraverseDirectory takes it from the filename it is already reading
  ("texts.txt" -> "texts"), the same rule ingest uses.
- CollectLikes has no file, so it reads the legacy-format `type` field that
  GrabLikes already requests with npf=false, mapped singular -> plural.
- AddPost/UpdatePost gained the plumbing to carry it. UpdatePost fills a
  missing type but never overwrites one, and its change-detection clause
  had to learn about PostType or the SET would be unreachable for a row
  whose content was already current.

PostTypes is the single source of truth: eight canonical names, and
anything else normalizes to null. Null is safe -- OutputMode skips those
rows -- while a stray value would have become a stray file. IngestMode and
TraverseDirectory now skip non-export .txt files outright, and the legacy
importer no longer passes a pre-column NULL straight back in.

For the rows already stored untyped, content is the only signal left, so
the migration infers from which columns they carry. Verified against a copy
of the live database: 1780 of 1790 typed, 10 left untyped for want of any
content at all, idempotent on a second pass. HasImage is deliberately not
consulted -- it is set on 12,420 of 19,828 known text posts.

Co-Authored-By: Claude Opus 5 <[email protected]>
2026-08-20 08:00:56 -05:00
9 changed files with 435 additions and 146 deletions
+1 -1
View File
@@ -31,7 +31,7 @@ dotnet run -- --test [blogname] [postID] # Test API for specific post
- `--test [blogname] [postID]`: Test API note collection
- `--posts`: Export post blogs to file
- `--blogs`: Export blog list to file
- `--collect [0|1] [datetime] [blogname]`: Collect notes for posts in DB. Optional `blogname` restricts the run to one blog (exact match), e.g. `--collect 1 zomb-eh`
- `--collect [0|1] [datetime] [blogname]`: Collect notes for posts in DB. Optional `blogname` restricts the run to one blog (exact match), e.g. `--collect 1 zomb-eh`. Add `--force` to ignore the periodic re-collect cooldown so already-collected posts are re-queued immediately (mode 1 only)
- `--blogsR`: Export reply blogs to file
- `--blogsO [start] [stop]`: Export blogs within range
+109 -100
View File
@@ -1,100 +1,109 @@
<?xml version="1.0" encoding="UTF-8"?><sqlb_project><db path="D:/NextCloud/C#/URLNotesGrabberCORE/URLNotesGrabberCORE/TL.db" readonly="0" foreign_keys="1" case_sensitive_like="0" temp_store="0" wal_autocheckpoint="1000" synchronous="2"/><attached/><window><main_tabs open="structure browser pragmas query" current="3"/></window><tab_structure><column_width id="0" width="300"/><column_width id="1" width="0"/><column_width id="2" width="100"/><column_width id="3" width="4486"/><column_width id="4" width="0"/><expanded_item id="0" parent="1"/><expanded_item id="1" parent="1"/><expanded_item id="2" parent="1"/><expanded_item id="3" parent="1"/></tab_structure><tab_browse><table title="Posts" custom_title="0" dock_id="4" table="4,5:mainPosts"/><dock_state state="000000ff00000000fd0000000100000002000005470000029afc0100000006fb000000160064006f0063006b00420072006f00770073006500310100000000000004a10000000000000000fb000000160064006f0063006b00420072006f00770073006500320100000000000004a10000000000000000fb000000160064006f0063006b00420072006f00770073006500330100000000000004a10000000000000000fb000000160064006f0063006b00420072006f00770073006500350100000000000005f40000000000000000fb000000160064006f0063006b00420072006f00770073006500340100000000000005470000011100fffffffb000000160064006f0063006b00420072006f00770073006500340100000000000005f40000000000000000000005470000000000000004000000040000000800000008fc00000000"/><default_encoding codec=""/><browse_table_settings><table schema="main" name="ApiKeyPoolMeta" show_row_id="0" encoding="" plot_x_axis="" unlock_view_pk="_rowid_" freeze_columns="0"><sort/><column_widths><column index="1" value="29"/><column index="2" value="64"/></column_widths><filter_values/><conditional_formats/><row_id_formats/><display_formats/><hidden_columns/><plot_y_axes/><global_filter/></table><table schema="main" name="Blogs" show_row_id="0" encoding="" plot_x_axis="" unlock_view_pk="_rowid_" freeze_columns="0"><sort><column index="7" mode="1"/></sort><column_widths><column index="1" value="257"/><column index="2" value="108"/><column index="3" value="63"/><column index="4" value="156"/><column index="5" value="60"/><column index="6" value="81"/><column index="7" value="85"/><column index="8" value="156"/><column index="9" value="156"/><column index="10" value="151"/><column index="11" value="125"/><column index="12" value="129"/></column_widths><filter_values><column index="4" value="1"/><column index="7" value="&gt;2026-05-27 17:22:36"/></filter_values><conditional_formats/><row_id_formats/><display_formats/><hidden_columns/><plot_y_axes/><global_filter/></table><table schema="main" name="Notes" show_row_id="0" encoding="" plot_x_axis="" unlock_view_pk="_rowid_" freeze_columns="0"><sort/><column_widths><column index="1" value="198"/><column index="2" value="144"/><column index="3" value="251"/><column index="4" value="84"/><column index="5" value="53"/><column index="6" value="300"/><column index="7" value="116"/><column index="8" value="152"/><column index="9" value="89"/><column index="10" value="63"/></column_widths><filter_values/><conditional_formats/><row_id_formats/><display_formats/><hidden_columns/><plot_y_axes/><global_filter/></table><table schema="main" name="Posts" show_row_id="0" encoding="" plot_x_axis="" unlock_view_pk="_rowid_" freeze_columns="0"><sort><column index="14" mode="1"/></sort><column_widths><column index="1" value="236"/><column index="2" value="144"/><column index="3" value="126"/><column index="4" value="32"/><column index="5" value="32"/><column index="6" value="32"/><column index="7" value="32"/><column index="8" value="32"/><column index="9" value="0"/><column index="10" value="0"/><column index="11" value="0"/><column index="12" value="243"/><column index="13" value="300"/><column index="14" value="53"/><column index="15" value="37351"/><column index="16" value="300"/><column index="17" value="41"/><column index="18" value="75"/><column index="19" value="96"/><column index="20" value="300"/><column index="21" value="96"/><column index="22" value="300"/><column index="23" value="300"/><column index="24" value="548"/><column index="25" value="60"/><column index="26" value="213"/><column index="27" value="532"/><column index="28" value="152"/><column index="29" value="152"/><column index="30" value="69"/><column index="31" value="63"/></column_widths><filter_values><column index="2" value="0"/><column index="1" value="734568371821084672"/></filter_values><conditional_formats/><row_id_formats/><display_formats/><hidden_columns><column index="9" value="1"/><column index="10" value="1"/><column index="11" value="1"/></hidden_columns><plot_y_axes/><global_filter/></table></browse_table_settings></tab_browse><tab_sql><sql name="SQL 1">UPDATE Posts
SET HasNotesGathered = 0
WHERE (BlogName, PostID) IN (
SELECT p.BlogName, p.PostID
FROM Posts p
WHERE p.HasNotesGathered = 1
AND P.notesGatheredDatetime &lt; 1774294520
AND EXISTS (
SELECT 1
FROM Notes n
WHERE n.PostID = p.PostID
AND n.RootBlogName = p.BlogName
--AND n.Type NOT IN ('reblog', 'reply')
)
ORDER BY P.PostDate ASC
--LIMIT 500
);</sql><sql name="Mark Blogs">select *
from Blogs
--update blogs set HasBeenOutput = 1
where HasBeenOutput = 0
AND
blogname in
(
'teaberrybee',
'reddevilgoddesstoo',
'waywardog13',
'wzjustbrowsing-blog',
'lewerta',
'nudenymph',
'caylachief'
)</sql><sql name="New Notes">select P.slug, N.replyText, n.RootBlogName, n.PostID, NoteBlogName || '.tumblr.com' as NoteBlogName, DatetimeCrawled, TimeStamp, type, n.RootBlogName || '.tumblr.com/post/' || n.postid, datetime(timestamp, 'unixepoch')
from Notes N inner join Posts P on p.PostID = n.PostID
where
DatetimeCrawled &gt; '2026-08-07 11:47:22' and type like 'r%'
and P.IsActive = 1
order by n.DatetimeCrawled</sql><sql name="Pull Blogs">SELECT distinct
'''' || blogname || ''',',
blogs.*
, blogname || '.tumblr.com'
FROM
Blogs
inner JOIN
Notes on notes.noteBlogName = blogs.BlogName
WHERE
HasBeenOutput = 0 and type = 'reblog'
order by
Notes.Type desc,
DateAdded desc
LIMIT 100;</sql><sql name="SQL 7">WITH ReplyCounts AS (
SELECT
NoteBlogName,
COUNT(DISTINCT replyText) AS DistinctReplyCount
FROM Notes
where replyText &lt;&gt; '.'
GROUP BY NoteBlogName
)
SELECT
n.RootBlogName || '.tumblr.com/post/' || n.PostID AS PostURL, postid,
n.NoteBlogName,
n.replyText,
c.DistinctReplyCount
FROM Notes n
JOIN ReplyCounts c ON n.NoteBlogName = c.NoteBlogName
where replyText &lt;&gt; '.' and type &lt;&gt; 'reply'
--AND N.NoteBlogName NOT IN ( 'roadblocker21', 'thesaddemon666', 'edwardabbeyhoffman', 'tattedsoldier20', 'zomb-eh', 'animalistic13', 'indken', 'maccloud1592',
--'moss-wizard', 'supertrucker12682', 'exploringthrupics', 'padeyepete' )
order by c.DistinctReplyCount desc, n.NoteBlogName, n.DateModified desc, replyText, RootBlogName, PostID</sql><sql name="Collect">WITH PostsWithCount AS ( SELECT P.BlogName, P.PostID, 1925013599 AS LatestNoteTimestamp, P.NotesGatheredDateTime, COUNT(P.PostID) OVER(PARTITION BY P.BlogName) AS CNT, P.HasNotesGathered, P.NotFound, P.PostDate FROM Posts P WHERE COALESCE(P.IsActive, 1) = 1 ), Unioned AS ( SELECT BlogName, PostID, LatestNoteTimestamp, NotesGatheredDateTime, CNT, PostDate FROM PostsWithCount WHERE NotFound = 0 AND HasNotesGathered = 0 UNION SELECT BlogName, PostID, LatestNoteTimestamp, NotesGatheredDateTime, CNT, PostDate FROM PostsWithCount WHERE BlogName = 'zomb-eh' AND NotFound = 0 AND NotesGatheredDateTime &lt; unixepoch('now', 'localtime', '-3 days') ) SELECT U.BlogName, U.PostID, U.LatestNoteTimestamp, U.NotesGatheredDateTime, U.CNT FROM Unioned U WHERE (U.NotesGatheredDateTime &lt; 1786134037 OR U.NotesGatheredDateTime IS NULL) ORDER BY U.NotesGatheredDateTime, U.PostDate DESC, U.BlogName, U.PostID;</sql><sql name="Del Posts">delete from posts where postid in
(
'741662499571728384',
178892849664,
178264721139,
177012868749,
169950081964,
755440787056099328
)</sql><sql name="notes NO post*">select *
-- delete
from notes
where postid not in (select distinct postid from posts where IsActive = 1)</sql><sql name="SQL 9">SELECT
*
FROM
POSTS P
WHERE
P.ByLikes = 1
AND
P.DateCreated &gt; '2026-05-26 17:47:32'
ORDER BY
P.DateCreated desc</sql><sql name="SQL 13">update posts set IsActive = 0 where blogname IN ( 'shoebiedoo', 'redheaded-girlygirl', 'xlittle-ghost' )</sql><sql name="SQL 14*">update Posts␍
set IsActive = 0␍
where postid in␍
(␍
␍
␍
'731937314675310592'␍
␍
␍
␍
)␍
␍
</sql><current_tab id="7"/></tab_sql></sqlb_project>
<?xml version="1.0" encoding="UTF-8"?><sqlb_project><db path="D:/NextCloud/C#/URLNotesGrabberCORE/URLNotesGrabberCORE/TL.db" readonly="0" foreign_keys="1" case_sensitive_like="0" temp_store="0" wal_autocheckpoint="1000" synchronous="2"/><attached/><window><main_tabs open="structure browser pragmas query" current="3"/></window><tab_structure><column_width id="0" width="300"/><column_width id="1" width="0"/><column_width id="2" width="100"/><column_width id="3" width="4486"/><column_width id="4" width="0"/><expanded_item id="0" parent="1"/><expanded_item id="1" parent="1"/><expanded_item id="2" parent="1"/><expanded_item id="3" parent="1"/></tab_structure><tab_browse><table title="Posts" custom_title="0" dock_id="4" table="4,5:mainPosts"/><dock_state state="000000ff00000000fd0000000100000002000005470000029afc0100000006fb000000160064006f0063006b00420072006f00770073006500310100000000000004a10000000000000000fb000000160064006f0063006b00420072006f00770073006500320100000000000004a10000000000000000fb000000160064006f0063006b00420072006f00770073006500330100000000000004a10000000000000000fb000000160064006f0063006b00420072006f00770073006500350100000000000005f40000000000000000fb000000160064006f0063006b00420072006f00770073006500340100000000000005470000011100fffffffb000000160064006f0063006b00420072006f00770073006500340100000000000005f40000000000000000000005470000000000000004000000040000000800000008fc00000000"/><default_encoding codec=""/><browse_table_settings><table schema="main" name="ApiKeyPoolMeta" show_row_id="0" encoding="" plot_x_axis="" unlock_view_pk="_rowid_" freeze_columns="0"><sort/><column_widths><column index="1" value="29"/><column index="2" value="64"/></column_widths><filter_values/><conditional_formats/><row_id_formats/><display_formats/><hidden_columns/><plot_y_axes/><global_filter/></table><table schema="main" name="Blogs" show_row_id="0" encoding="" plot_x_axis="" unlock_view_pk="_rowid_" freeze_columns="0"><sort><column index="7" mode="1"/></sort><column_widths><column index="1" value="257"/><column index="2" value="108"/><column index="3" value="63"/><column index="4" value="156"/><column index="5" value="60"/><column index="6" value="81"/><column index="7" value="85"/><column index="8" value="156"/><column index="9" value="156"/><column index="10" value="151"/><column index="11" value="125"/><column index="12" value="129"/></column_widths><filter_values><column index="4" value="1"/><column index="7" value="&gt;2026-05-27 17:22:36"/></filter_values><conditional_formats/><row_id_formats/><display_formats/><hidden_columns/><plot_y_axes/><global_filter/></table><table schema="main" name="Notes" show_row_id="0" encoding="" plot_x_axis="" unlock_view_pk="_rowid_" freeze_columns="0"><sort/><column_widths><column index="1" value="198"/><column index="2" value="144"/><column index="3" value="251"/><column index="4" value="84"/><column index="5" value="53"/><column index="6" value="300"/><column index="7" value="116"/><column index="8" value="152"/><column index="9" value="89"/><column index="10" value="63"/></column_widths><filter_values/><conditional_formats/><row_id_formats/><display_formats/><hidden_columns/><plot_y_axes/><global_filter/></table><table schema="main" name="Posts" show_row_id="0" encoding="" plot_x_axis="" unlock_view_pk="_rowid_" freeze_columns="0"><sort><column index="14" mode="1"/></sort><column_widths><column index="1" value="236"/><column index="2" value="144"/><column index="3" value="126"/><column index="4" value="32"/><column index="5" value="32"/><column index="6" value="32"/><column index="7" value="32"/><column index="8" value="32"/><column index="9" value="0"/><column index="10" value="0"/><column index="11" value="0"/><column index="12" value="243"/><column index="13" value="300"/><column index="14" value="53"/><column index="15" value="37351"/><column index="16" value="300"/><column index="17" value="41"/><column index="18" value="75"/><column index="19" value="96"/><column index="20" value="300"/><column index="21" value="96"/><column index="22" value="300"/><column index="23" value="300"/><column index="24" value="548"/><column index="25" value="60"/><column index="26" value="213"/><column index="27" value="532"/><column index="28" value="152"/><column index="29" value="152"/><column index="30" value="69"/><column index="31" value="63"/></column_widths><filter_values><column index="2" value="0"/><column index="1" value="734568371821084672"/></filter_values><conditional_formats/><row_id_formats/><display_formats/><hidden_columns><column index="9" value="1"/><column index="10" value="1"/><column index="11" value="1"/></hidden_columns><plot_y_axes/><global_filter/></table></browse_table_settings></tab_browse><tab_sql><sql name="SQL 1">UPDATE Posts
SET HasNotesGathered = 0
WHERE (BlogName, PostID) IN (
SELECT p.BlogName, p.PostID
FROM Posts p
WHERE p.HasNotesGathered = 1
AND P.notesGatheredDatetime &lt; 1774294520
AND EXISTS (
SELECT 1
FROM Notes n
WHERE n.PostID = p.PostID
AND n.RootBlogId = (SELECT BlogId FROM BlogNames WHERE BlogName = p.BlogName)
--AND n.TypeId NOT IN (SELECT TypeId FROM NoteTypes WHERE Type IN ('reblog', 'reply'))
)
ORDER BY P.PostDate ASC
--LIMIT 500
);</sql><sql name="Mark Blogs">select *
from Blogs
--update blogs set HasBeenOutput = 1
where HasBeenOutput = 0
AND
blogname in
(
'teaberrybee',
'reddevilgoddesstoo',
'waywardog13',
'wzjustbrowsing-blog',
'lewerta',
'nudenymph',
'caylachief'
)</sql><sql name="New Notes">select P.slug, N.replyText, rbn.BlogName as RootBlogName, n.PostID, nbn.BlogName || '.tumblr.com' as NoteBlogName, DatetimeCrawled, TimeStamp, nt.Type, rbn.BlogName || '.tumblr.com/post/' || n.postid, datetime(timestamp, 'unixepoch')
from Notes N
inner join Posts P on p.PostID = n.PostID
inner join BlogNames rbn on rbn.BlogId = n.RootBlogId
inner join BlogNames nbn on nbn.BlogId = n.NoteBlogId
inner join NoteTypes nt on nt.TypeId = n.TypeId
where
DatetimeCrawled &gt; '2026-08-07 11:47:22' and nt.Type like 'r%'
and P.IsActive = 1
order by n.DatetimeCrawled</sql><sql name="Pull Blogs">SELECT distinct
'''' || blogname || ''',',
blogs.*
, blogname || '.tumblr.com'
FROM
Blogs
inner JOIN
Notes on notes.noteBlogId = blogs.BlogId
inner JOIN
NoteTypes on NoteTypes.TypeId = Notes.TypeId
WHERE
HasBeenOutput = 0 and NoteTypes.Type = 'reblog'
order by
NoteTypes.Type desc,
DateAdded desc
LIMIT 100;</sql><sql name="SQL 7">WITH ReplyCounts AS (
SELECT
NoteBlogId,
COUNT(DISTINCT replyText) AS DistinctReplyCount
FROM Notes
where replyText &lt;&gt; '.'
GROUP BY NoteBlogId
)
SELECT
rbn.BlogName || '.tumblr.com/post/' || n.PostID AS PostURL, postid,
nbn.BlogName AS NoteBlogName,
n.replyText,
c.DistinctReplyCount
FROM Notes n
JOIN ReplyCounts c ON n.NoteBlogId = c.NoteBlogId
JOIN BlogNames rbn ON rbn.BlogId = n.RootBlogId
JOIN BlogNames nbn ON nbn.BlogId = n.NoteBlogId
JOIN NoteTypes t ON t.TypeId = n.TypeId
where replyText &lt;&gt; '.' and t.Type &lt;&gt; 'reply'
--AND nbn.BlogName NOT IN ( 'roadblocker21', 'thesaddemon666', 'edwardabbeyhoffman', 'tattedsoldier20', 'zomb-eh', 'animalistic13', 'indken', 'maccloud1592',
--'moss-wizard', 'supertrucker12682', 'exploringthrupics', 'padeyepete' )
order by c.DistinctReplyCount desc, nbn.BlogName, n.DateModified desc, replyText, rbn.BlogName, PostID</sql><sql name="Collect">WITH PostsWithCount AS ( SELECT P.BlogName, P.PostID, 1925013599 AS LatestNoteTimestamp, P.NotesGatheredDateTime, COUNT(P.PostID) OVER(PARTITION BY P.BlogName) AS CNT, P.HasNotesGathered, P.NotFound, P.PostDate FROM Posts P WHERE COALESCE(P.IsActive, 1) = 1 ), Unioned AS ( SELECT BlogName, PostID, LatestNoteTimestamp, NotesGatheredDateTime, CNT, PostDate FROM PostsWithCount WHERE NotFound = 0 AND HasNotesGathered = 0 UNION SELECT BlogName, PostID, LatestNoteTimestamp, NotesGatheredDateTime, CNT, PostDate FROM PostsWithCount WHERE BlogName = 'zomb-eh' AND NotFound = 0 AND NotesGatheredDateTime &lt; unixepoch('now', 'localtime', '-3 days') ) SELECT U.BlogName, U.PostID, U.LatestNoteTimestamp, U.NotesGatheredDateTime, U.CNT FROM Unioned U WHERE (U.NotesGatheredDateTime &lt; 1786134037 OR U.NotesGatheredDateTime IS NULL) ORDER BY U.NotesGatheredDateTime, U.PostDate DESC, U.BlogName, U.PostID;</sql><sql name="Del Posts">delete from posts where postid in
(
'741662499571728384',
178892849664,
178264721139,
177012868749,
169950081964,
755440787056099328
)</sql><sql name="notes NO post*">select *
-- delete
from notes
where postid not in (select distinct postid from posts where IsActive = 1)</sql><sql name="SQL 9">SELECT
*
FROM
POSTS P
WHERE
P.ByLikes = 1
AND
P.DateCreated &gt; '2026-05-26 17:47:32'
ORDER BY
P.DateCreated desc</sql><sql name="SQL 13">update posts set IsActive = 0 where blogname IN ( 'shoebiedoo', 'redheaded-girlygirl', 'xlittle-ghost' )</sql><sql name="SQL 14*">update Posts
set IsActive = 0
where postid in
(
'731937314675310592'
)
</sql><current_tab id="7"/></tab_sql></sqlb_project>
+30 -24
View File
@@ -1,24 +1,30 @@
<?xml version="1.0" encoding="UTF-8"?><sqlb_project><db path="" readonly="1" foreign_keys="" case_sensitive_like="" temp_store="" wal_autocheckpoint="" synchronous=""/><attached/><window><main_tabs open="structure browser pragmas query" current="3"/></window><tab_structure><column_width id="0" width="300"/><column_width id="1" width="0"/><column_width id="2" width="100"/><column_width id="3" width="3571"/><column_width id="4" width="0"/></tab_structure><tab_browse><table title="." custom_title="0" dock_id="4" table="0,0:"/><dock_state state="000000ff00000000fd0000000100000002000005f40000030ffc0100000002fb000000160064006f0063006b00420072006f00770073006500310100000000000005f40000000000000000fb000000160064006f0063006b00420072006f00770073006500340100000000ffffffff0000011700ffffff000005f40000000000000004000000040000000800000008fc00000000"/><default_encoding codec=""/><browse_table_settings/></tab_browse><tab_sql><sql name="SQL 1">SELECT
BlogName || '.tumblr.com/post/' || postID,
datetime(NotesGatheredDateTime, 'unixepoch'), *
FROM
Posts
WHERE
NotesGatheredDateTime &lt;&gt; 0
ORDER BY
postdate desc</sql><sql name="SQL 2*">SELECT
datetime(TimeStamp, 'unixepoch'),
RootBlogName || '.tumblr.com/post/' || N.postid,
*,
NoteBlogName || '.tumblr.com'
FROM
Notes N␍
inner JOIN␍
Posts P on P.PostID = N.PostID and P.BlogName = N.RootBlogName
WHERE RootBlogName NOT IN ('xlittle-ghost', 'glimmerin-darlin', 'vvenus-child')␍
and type like 'r%'␍
and RootBlogName = 'zomb-eh'␍
and P.HasImage = 1
ORDER BY
TimeStamp desc</sql><current_tab id="1"/></tab_sql></sqlb_project>
<?xml version="1.0" encoding="UTF-8"?><sqlb_project><db path="" readonly="1" foreign_keys="" case_sensitive_like="" temp_store="" wal_autocheckpoint="" synchronous=""/><attached/><window><main_tabs open="structure browser pragmas query" current="3"/></window><tab_structure><column_width id="0" width="300"/><column_width id="1" width="0"/><column_width id="2" width="100"/><column_width id="3" width="3571"/><column_width id="4" width="0"/></tab_structure><tab_browse><table title="." custom_title="0" dock_id="4" table="0,0:"/><dock_state state="000000ff00000000fd0000000100000002000005f40000030ffc0100000002fb000000160064006f0063006b00420072006f00770073006500310100000000000005f40000000000000000fb000000160064006f0063006b00420072006f00770073006500340100000000ffffffff0000011700ffffff000005f40000000000000004000000040000000800000008fc00000000"/><default_encoding codec=""/><browse_table_settings/></tab_browse><tab_sql><sql name="SQL 1">SELECT
BlogName || '.tumblr.com/post/' || postID,
datetime(NotesGatheredDateTime, 'unixepoch'), *
FROM
Posts
WHERE
NotesGatheredDateTime &lt;&gt; 0
ORDER BY
postdate desc</sql><sql name="SQL 2*">SELECT
datetime(TimeStamp, 'unixepoch'),
rbn.BlogName || '.tumblr.com/post/' || N.postid,
*,
nbn.BlogName || '.tumblr.com'
FROM
Notes N
inner JOIN
BlogNames rbn on rbn.BlogId = N.RootBlogId
inner JOIN
BlogNames nbn on nbn.BlogId = N.NoteBlogId
inner JOIN
NoteTypes t on t.TypeId = N.TypeId
inner JOIN
Posts P on P.PostID = N.PostID and P.BlogName = rbn.BlogName
WHERE rbn.BlogName NOT IN ('xlittle-ghost', 'glimmerin-darlin', 'vvenus-child')
and t.Type like 'r%'
and rbn.BlogName = 'zomb-eh'
and P.HasImage = 1
ORDER BY
TimeStamp desc</sql><current_tab id="1"/></tab_sql></sqlb_project>
+95 -8
View File
@@ -590,12 +590,16 @@ namespace URLNotesGrabberCORE
}
}
public static void AddPost(string blogName, long postID, string reblogURL, string postDate, string postURL, string slug, string reblogKey, string reblogName, string summary, string quote, string body, string tags, string link, string photoURL, string photoCaption, string downloadedFiles, string audioCaption, string question, string answer, string title, bool hasImage, bool byLikes = false, string? DBPath = null, string? rootBlogName = null, string? rootURL = null)
// postType: a canonical PostTypes name, or null when the caller has no trustworthy type.
// Null is stored as NULL rather than guessed at -- OutputMode skips untyped rows, so a
// null costs one export line, whereas a wrong value would create a wrongly named file.
public static void AddPost(string blogName, long postID, string reblogURL, string postDate, string postURL, string slug, string reblogKey, string reblogName, string summary, string quote, string body, string tags, string link, string photoURL, string photoCaption, string downloadedFiles, string audioCaption, string question, string answer, string title, bool hasImage, bool byLikes = false, string? DBPath = null, string? rootBlogName = null, string? rootURL = null, string? postType = null)
{
DBPath ??= GetDefaultDbPath();
postType = PostTypes.Normalize(postType);
try { AddBlog(blogName, byLikes, DBPath); } catch { }
try { UpdatePostSetDate(blogName, postID, postDate, DBPath); } catch { }
try { UpdatePost(blogName, postID, reblogURL, postDate, postURL, slug, reblogKey, reblogName, summary, quote, body, tags, link, photoURL, photoCaption, downloadedFiles, audioCaption, question, answer, title, hasImage, byLikes, DBPath, rootBlogName, rootURL); } catch { }
try { UpdatePost(blogName, postID, reblogURL, postDate, postURL, slug, reblogKey, reblogName, summary, quote, body, tags, link, photoURL, photoCaption, downloadedFiles, audioCaption, question, answer, title, hasImage, byLikes, DBPath, rootBlogName, rootURL, postType); } catch { }
SQLiteConnection connection;
bool ownsConnection;
@@ -637,9 +641,10 @@ namespace URLNotesGrabberCORE
RootBlogName,
RootURL,
HasImage,
ByLikes
ByLikes,
PostType
) VALUES (" +
Q(blogName) + ", " + postID + ", " + Q(reblogURL) + ", " + Q(postDate) + ", " + Q(postURL) + ", " + Q(slug) + ", " + Q(reblogKey) + ", " + Q(reblogName) + ", " + Q(summary) + ", " + Q(quote) + ", " + Q(body) + ", " + Q(tags) + ", " + Q(link) + ", " + Q(photoURL) + ", " + Q(photoCaption) + ", " + Q(downloadedFiles) + ", " + Q(audioCaption) + ", " + Q(question) + ", " + Q(answer) + ", " + Q(title) + ", " + Q(DateTime.Now.ToString("yyyy-MM-dd HH:mm:ss")) + ", " + Q(DateTime.Now.ToString("yyyy-MM-dd HH:mm:ss")) + ", " + Q(rootBlogName ?? ".") + ", " + Q(rootURL ?? ".") + ", " + (hasImage ? 1 : 0) + ", " + (byLikes ? 1 : 0) + ")";
Q(blogName) + ", " + postID + ", " + Q(reblogURL) + ", " + Q(postDate) + ", " + Q(postURL) + ", " + Q(slug) + ", " + Q(reblogKey) + ", " + Q(reblogName) + ", " + Q(summary) + ", " + Q(quote) + ", " + Q(body) + ", " + Q(tags) + ", " + Q(link) + ", " + Q(photoURL) + ", " + Q(photoCaption) + ", " + Q(downloadedFiles) + ", " + Q(audioCaption) + ", " + Q(question) + ", " + Q(answer) + ", " + Q(title) + ", " + Q(DateTime.Now.ToString("yyyy-MM-dd HH:mm:ss")) + ", " + Q(DateTime.Now.ToString("yyyy-MM-dd HH:mm:ss")) + ", " + Q(rootBlogName ?? ".") + ", " + Q(rootURL ?? ".") + ", " + (hasImage ? 1 : 0) + ", " + (byLikes ? 1 : 0) + ", " + (postType == null ? "NULL" : Q(postType)) + ")";
SQLiteCommand command = new SQLiteCommand(sql, connection);
int rowsInserted = 0;
@@ -853,9 +858,10 @@ namespace URLNotesGrabberCORE
///
/// </summary>
/// <param name="withoutNotesOnly"></param>
/// <param name="ignoreRefreshCooldown">Drops the age gate on the periodic re-queue branch (--force).</param>
/// <param name="DBPath"></param>
/// <returns>blogName, postID, lastNoteTimestamp, notesGatheredTimestamp</returns>
public static List<Tuple<string, long, long, long>> GetPosts(bool withoutNotesOnly = false, DateTime? beforeDate = null, string? blogName = null, string? DBPath = null)
public static List<Tuple<string, long, long, long>> GetPosts(bool withoutNotesOnly = false, DateTime? beforeDate = null, string? blogName = null, bool ignoreRefreshCooldown = false, string? DBPath = null)
{
DBPath ??= GetDefaultDbPath();
using SQLiteConnection connection = new SQLiteConnection("Data Source=" + DBPath);
@@ -890,6 +896,13 @@ namespace URLNotesGrabberCORE
// scoped, so it contributes its rows when the filter names zomb-eh and nothing otherwise.
// That keeps a filtered worklist a strict subset of the unfiltered one -- "--collect 1 X"
// returns exactly the rows "--collect 1" would have returned for X.
//
// --force drops the age gate only. NotFound = 0 and the IsActive/blog scoping above still
// apply: the flag is "re-collect early", not "collect rows every other path excludes".
string refreshCooldownClause = ignoreRefreshCooldown
? string.Empty
: " AND NotesGatheredDateTime < unixepoch('now', 'localtime', '-3 days')" + Environment.NewLine;
string refreshBranch =
"" + Environment.NewLine +
" UNION " + Environment.NewLine +
@@ -904,7 +917,7 @@ namespace URLNotesGrabberCORE
" FROM PostsWithCount" + Environment.NewLine +
" WHERE BlogName = 'zomb-eh'" + Environment.NewLine +
" AND NotFound = 0" + Environment.NewLine +
" AND NotesGatheredDateTime < unixepoch('now', 'localtime', '-3 days')" + Environment.NewLine;
refreshCooldownClause;
sql = "WITH PostsWithCount AS" + Environment.NewLine +
"(" + Environment.NewLine +
@@ -1679,9 +1692,10 @@ namespace URLNotesGrabberCORE
return false;
}
public static void UpdatePost(string blogName, long postID, string reblogURL, string postDate, string postURL, string slug, string reblogKey, string reblogName, string summary, string quote, string body, string tags, string link, string photoURL, string photoCaption, string downloadedFiles, string audioCaption, string question, string answer, string title, bool hasImage, bool byLikes = false, string? DBPath = null, string? rootBlogName = null, string? rootURL = null)
public static void UpdatePost(string blogName, long postID, string reblogURL, string postDate, string postURL, string slug, string reblogKey, string reblogName, string summary, string quote, string body, string tags, string link, string photoURL, string photoCaption, string downloadedFiles, string audioCaption, string question, string answer, string title, bool hasImage, bool byLikes = false, string? DBPath = null, string? rootBlogName = null, string? rootURL = null, string? postType = null)
{
DBPath ??= GetDefaultDbPath();
postType = PostTypes.Normalize(postType);
SQLiteConnection connection;
bool ownsConnection;
@@ -1729,6 +1743,10 @@ namespace URLNotesGrabberCORE
sql += "RootBlogName = CASE WHEN @rootBlogName IS NULL OR @rootBlogName = '' OR @rootBlogName = '.' THEN RootBlogName ELSE @rootBlogName END, ";
sql += "RootURL = CASE WHEN @rootURL IS NULL OR @rootURL = '' OR @rootURL = '.' THEN RootURL ELSE @rootURL END, ";
sql += "hasImage = @hasImage, ";
// Fill in a missing type, never overwrite one. A type derived by --ingest from a
// real export filename is authoritative; this path's type is only as good as the
// folder it was crawled from, so it must not win over an existing value.
sql += "PostType = IFNULL(PostType, @postType), ";
sql += "ByLikes = MAX(IFNULL(ByLikes, 0), @byLikes) ";
sql += " WHERE BlogName = @BlogName AND PostID = @PostID AND (";
sql += "(@postDate <> '.' AND IFNULL(postDate, '') <> @postDate) OR ";
@@ -1752,7 +1770,10 @@ namespace URLNotesGrabberCORE
sql += "IFNULL(hasImage, 0) <> @hasImage OR ";
sql += "(@byLikes = 1 AND IFNULL(ByLikes, 0) = 0) OR ";
sql += "((@rootBlogName IS NOT NULL AND @rootBlogName <> '' AND @rootBlogName <> '.') AND IFNULL(RootBlogName, '') <> @rootBlogName) OR ";
sql += "((@rootURL IS NOT NULL AND @rootURL <> '' AND @rootURL <> '.') AND IFNULL(RootURL, '') <> @rootURL)";
sql += "((@rootURL IS NOT NULL AND @rootURL <> '' AND @rootURL <> '.') AND IFNULL(RootURL, '') <> @rootURL) OR ";
// Without this the SET above is unreachable for a row whose content is already
// current: the UPDATE would not fire, and the type would stay NULL forever.
sql += "(PostType IS NULL AND @postType IS NOT NULL)";
sql += ")";
using (SQLiteCommand command = new SQLiteCommand(sql, connection))
@@ -1780,6 +1801,7 @@ namespace URLNotesGrabberCORE
command.Parameters.AddWithValue("@rootURL", string.IsNullOrWhiteSpace(rootURL) ? (object)DBNull.Value : rootURL);
command.Parameters.AddWithValue("@hasImage", hasImage ? 1 : 0);
command.Parameters.AddWithValue("@byLikes", byLikes ? 1 : 0);
command.Parameters.AddWithValue("@postType", (object?)postType ?? DBNull.Value);
command.Parameters.AddWithValue("@BlogName", blogName);
command.Parameters.AddWithValue("@PostID", postID);
@@ -2089,6 +2111,8 @@ namespace URLNotesGrabberCORE
cmd.ExecuteNonQuery();
Console.WriteLine("[Migration] Added PostType column to Posts table");
}
BackfillMissingPostTypes(connection);
}
catch (Exception ex)
{
@@ -2096,6 +2120,65 @@ namespace URLNotesGrabberCORE
}
}
/// <summary>
/// Types rows that carry no PostType, inferring it from which content columns they hold.
///
/// These are posts harvested from notes and likes rather than read out of a TumblThree
/// export, so no filename ever described them and --ingest can never reach them: it only
/// types a post it meets inside a real .txt. Content is the only signal they have.
///
/// Runs on every migration pass and is idempotent -- it only touches PostType IS NULL,
/// so a row typed once is never revisited. Rows whose columns give no signal at all stay
/// NULL and are skipped by OutputMode.
///
/// Mirrors PostTypes.InferFromContent; the two must agree. Notably HasImage is not
/// consulted, because most text posts carry it.
/// </summary>
private static void BackfillMissingPostTypes(SQLiteConnection connection)
{
const string set = @"
UPDATE Posts SET PostType = CASE
WHEN Has(Question) AND Has(Answer) THEN 'answers'
WHEN Has(Quote) THEN 'quotes'
WHEN Has(Link) THEN 'links'
WHEN Has(AudioCaption) THEN 'audios'
WHEN Has(Body) THEN 'texts'
WHEN Has(PhotoURL) OR Has(PhotoCaption) THEN 'images'
ELSE NULL END
WHERE PostType IS NULL";
// SQLite has no user-defined predicate here, so expand the "field supplied" test
// ("." is the not-supplied sentinel used throughout the export format) inline.
string sql = System.Text.RegularExpressions.Regex.Replace(
set, @"Has\((\w+)\)", "TRIM(IFNULL($1, '')) NOT IN ('', '.')");
try
{
long before;
using (var count = new SQLiteCommand("SELECT COUNT(*) FROM Posts WHERE PostType IS NULL", connection))
before = Convert.ToInt64(count.ExecuteScalar());
if (before == 0) return;
int changed;
using (var cmd = new SQLiteCommand(sql, connection))
changed = cmd.ExecuteNonQuery();
long after;
using (var count = new SQLiteCommand("SELECT COUNT(*) FROM Posts WHERE PostType IS NULL", connection))
after = Convert.ToInt64(count.ExecuteScalar());
if (changed > 0 || after != before)
Console.WriteLine($"[Migration] Backfilled PostType for {before - after} post(s); {after} still untyped (no content signal).");
}
catch (Exception ex)
{
// A failed backfill must not stop the run: untyped rows are skipped on export,
// which is inconvenient, not corrupting.
Console.WriteLine($"[Migration] PostType backfill failed: {ex.Message}");
}
}
// INSERT-or-UPDATE for a post arriving from a Tumblr text-file export.
// On collision, only content columns + PostType + DateModified are updated;
// engagement columns (ByLikes, RootBlogName, RootURL, HasNotesGathered, NotFound,
@@ -2126,6 +2209,10 @@ namespace URLNotesGrabberCORE
string? DBPath = null)
{
DBPath ??= GetDefaultDbPath();
// Central guarantee: whatever a caller believes, only a canonical type reaches the
// column. PostType is used as an output filename, so this is the invariant that keeps
// a stray value from becoming a stray file.
postType = PostTypes.Normalize(postType);
try { AddBlog(blogName, false, DBPath); } catch { }
SQLiteConnection connection;
+19 -1
View File
@@ -85,7 +85,25 @@ namespace URLNotesGrabberCORE
{
string rawBlogName = Path.GetFileName(Path.GetDirectoryName(file) ?? "unknown");
string blogName = Regex.Replace(rawBlogName, @"_\d+$", "");
string postType = Path.GetFileNameWithoutExtension(file);
// The filename becomes the row's PostType, and PostType later becomes an
// output filename -- so an unrecognized name here would mint a new type and
// a new file from any stray .txt that happens to sit in the tree. Only the
// eight real export files are ingestable.
//
// This is also what breaks the Unknown.txt cycle: OutputMode used to write
// untyped rows to Unknown.txt, and this scan would read it straight back
// and stamp those rows with the literal type "Unknown", making the file
// regenerate itself forever.
string? resolvedPostType = PostTypes.FromFileName(file);
if (resolvedPostType == null)
{
filesSkipped++;
continue;
}
// Non-nullable from here so the local Flush() below stays warning-clean:
// nullable flow analysis does not reach into local functions.
string postType = resolvedPostType;
if (targetBlog != null && !string.Equals(blogName, targetBlog, StringComparison.OrdinalIgnoreCase))
{
+6 -1
View File
@@ -117,7 +117,12 @@ namespace URLNotesGrabberCORE
question: reader.IsDBNull(18) ? null : reader.GetString(18),
answer: reader.IsDBNull(19) ? null : reader.GetString(19),
title: reader.IsDBNull(20) ? null : reader.GetString(20),
postType: reader.IsDBNull(21) ? null : reader.GetString(21),
// A legacy Posts.db predating the PostType column hands back NULL
// here, and on the INSERT branch that NULL is stored -- reseeding
// exactly the untyped rows the backfill exists to clear. Normalize
// so an unrecognized legacy value cannot become a filename either;
// the backfill types whatever comes through as null.
postType: PostTypes.Normalize(reader.IsDBNull(21) ? null : reader.GetString(21)),
hasImage: hasImage);
postsUpserted++;
if (postsUpserted % 500 == 0)
+12 -2
View File
@@ -65,10 +65,20 @@ namespace URLNotesGrabberCORE
var posts = DataAccess.GetAllPostsForBlog(blogName);
Console.WriteLine($" Found {posts.Count} post(s) for this blog.");
var grouped = posts.GroupBy(p => p.PostType ?? "Unknown");
// A post's type becomes a filename, so only a recognized type may be written. The
// old `PostType ?? "Unknown"` invented Unknown.txt for untyped rows, which --ingest
// then read back as a type named "Unknown" -- the two regenerated each other.
// Untyped rows are skipped instead: after the backfill these are only rows with no
// content signal at all, so nothing meaningful is lost, and nothing is invented.
var typed = posts.Where(p => PostTypes.Normalize(p.PostType) != null).ToList();
int untyped = posts.Count - typed.Count;
if (untyped > 0)
Console.WriteLine($" Skipping {untyped} post(s) with no recognized PostType.");
var grouped = typed.GroupBy(p => PostTypes.Normalize(p.PostType)!);
foreach (var typeGroup in grouped)
{
string postType = typeGroup.Key ?? "Unknown";
string postType = typeGroup.Key;
string outputFilePath = Path.Combine(folder, $"{postType}.txt");
var ordered = typeGroup.OrderBy(p => p.Date).ToList();
Console.WriteLine($" Writing {ordered.Count} post(s) to {postType}.txt");
+123
View File
@@ -0,0 +1,123 @@
using System;
using System.Collections.Generic;
using System.IO;
namespace URLNotesGrabberCORE
{
/// <summary>
/// The single source of truth for Posts.PostType values.
///
/// PostType exists so --output can write one .txt per type. Because the type becomes a
/// *filename*, an unvalidated value is not a cosmetic problem: it creates a file. That is
/// how "Unknown.txt" came about -- OutputMode used `PostType ?? "Unknown"` as a filename,
/// --ingest then read that file straight back and derived the literal type "Unknown" from
/// its name, and the pair would have kept regenerating each other indefinitely.
///
/// So every path that produces a type routes through <see cref="Normalize"/>, which admits
/// only the eight known names and returns null for anything else. A null type is safe:
/// OutputMode skips those rows rather than inventing a file for them.
/// </summary>
public static class PostTypes
{
// The canonical set. These are exactly the TumblThree .txt basenames, which is what
// makes an ingested filename usable as a type without translation.
public const string Texts = "texts";
public const string Answers = "answers";
public const string Quotes = "quotes";
public const string Links = "links";
public const string Conversations = "conversations";
public const string Images = "images";
public const string Videos = "videos";
public const string Audios = "audios";
private static readonly HashSet<string> Known = new HashSet<string>(
new[] { Texts, Answers, Quotes, Links, Conversations, Images, Videos, Audios },
StringComparer.OrdinalIgnoreCase);
// Tumblr's legacy post format (npf=false) names types in the singular. The likes API is
// the one source that reports a type directly rather than via a filename, so it is the
// only place this mapping is needed.
private static readonly Dictionary<string, string> ApiTypeMap = new Dictionary<string, string>(StringComparer.OrdinalIgnoreCase)
{
["text"] = Texts,
["photo"] = Images,
["quote"] = Quotes,
["link"] = Links,
["chat"] = Conversations,
["answer"] = Answers,
["audio"] = Audios,
["video"] = Videos,
};
/// <summary>
/// Returns the canonical type name, or null if the value is not one of the eight.
/// Returning null rather than passing the value through is the whole point: an
/// unrecognized string must never reach a filename.
/// </summary>
public static string? Normalize(string? candidate)
{
if (string.IsNullOrWhiteSpace(candidate)) return null;
string trimmed = candidate.Trim();
return Known.TryGetValue(trimmed, out string? canonical) ? canonical : null;
}
/// <summary>
/// Type for a post read out of a TumblThree export file, taken from the filename
/// ("texts.txt" -> "texts"). Anything else in the folder -- README.txt, a stray
/// triage file, or a previously written Unknown.txt -- normalizes to null and is
/// rejected by the caller.
/// </summary>
public static string? FromFileName(string? path)
{
if (string.IsNullOrWhiteSpace(path)) return null;
return Normalize(Path.GetFileNameWithoutExtension(path));
}
/// <summary>
/// Type for a post from the likes API, whose legacy-format `type` field is singular.
/// Null when the field is absent or unrecognized -- the access is dynamic, so a missing
/// field yields null at runtime rather than failing to compile.
/// </summary>
public static string? FromApiType(string? apiType)
{
if (string.IsNullOrWhiteSpace(apiType)) return null;
return ApiTypeMap.TryGetValue(apiType.Trim(), out string? mapped) ? mapped : null;
}
/// <summary>
/// Last-resort type inferred from which content columns a row actually carries. Used
/// only to backfill rows written before any type was recorded; a filename or an API
/// type is always preferred over this.
///
/// The order matters and is derived from the already-typed rows, where the column
/// signatures are effectively disjoint: answers carry Question+Answer and no Body,
/// images carry photo columns and no Body, texts carry Body and no photo columns.
///
/// HasImage is deliberately NOT consulted: it is set on 12,420 of 19,828 known text
/// posts, so it says nothing about the post's type.
///
/// conversations cannot be separated from texts this way -- both carry only Body -- so
/// a chat post with no other signal is labelled texts. A later --ingest that meets the
/// post in a real conversations.txt corrects it.
/// </summary>
public static string? InferFromContent(string? question, string? answer, string? quote,
string? link, string? audioCaption, string? body, string? photoUrl, string? photoCaption)
{
if (HasValue(question) && HasValue(answer)) return Answers;
if (HasValue(quote)) return Quotes;
if (HasValue(link)) return Links;
if (HasValue(audioCaption)) return Audios;
if (HasValue(body)) return Texts;
if (HasValue(photoUrl) || HasValue(photoCaption)) return Images;
return null;
}
// "." is the codebase-wide "field not supplied" sentinel in export records, so it
// counts as absent here just as it does in UpdatePost's CASE guards.
private static bool HasValue(string? value)
{
if (string.IsNullOrWhiteSpace(value)) return false;
return value.Trim() != ".";
}
}
}
+40 -9
View File
@@ -148,6 +148,12 @@ namespace URLNotesGrabberCORE
if (args.Length == 0) //Traverse folder structure to add posts and thus blogs to DB
{
int postsAdded = 0;
// This is the mode that actually gets run day to day, so the schema migration and
// the PostType backfill have to happen here too. They used to hang off --ingest,
// --output and friends only, which meant the untyped rows this traversal creates
// could sit unrepaired indefinitely while the one command everyone runs skipped
// the fix entirely. Idempotent, so paying it on every run costs nothing.
DataAccess.EnsureTTFileHelperColumnsExist();
try
{
DataAccess.EnableImportModePragmas();
@@ -316,7 +322,12 @@ namespace URLNotesGrabberCORE
managedCollectRun = true;
}
exitCode = CollectNotes(settings.GetValue<string>("PathOutput"), withoutNotesOnly, beforeDate, managedCollectRun, collectBlogName).GetAwaiter().GetResult();
if (forceIgnoreCooldown)
Console.WriteLine(withoutNotesOnly
? "--force: ignoring the periodic re-collect cooldown; already-collected posts in scope are re-queued now"
: "--force: no effect in mode 0 - a full re-check already re-collects every post");
exitCode = CollectNotes(settings.GetValue<string>("PathOutput"), withoutNotesOnly, beforeDate, managedCollectRun, collectBlogName, forceIgnoreCooldown).GetAwaiter().GetResult();
break;
case "--blogsR": //collect notes from all posts
@@ -448,7 +459,7 @@ namespace URLNotesGrabberCORE
Console.WriteLine("--blogs\t For each Blog in DB, write blogname to file");
Console.WriteLine("--collect [0|1] [datetime] [blogname]\t Collect Notes from API. 1=only posts without notes. 0=full re-check of all posts: a single resumable pass (interrupt & relaunch to resume; stops when complete, retrigger for a new pass). Optional datetime overrides the cutoff and runs as a one-off (bypasses resume tracking). Optional blogname restricts the run to that blog (exact, case-sensitive match) and also runs as a one-off; e.g. \"--collect 1 zomb-eh\". datetime and blogname may be given in either order - use --blog=name if a blog name would otherwise parse as a date.");
Console.WriteLine("--collect [0|1] [datetime] [blogname]\t Collect Notes from API. 1=only posts without notes. 0=full re-check of all posts: a single resumable pass (interrupt & relaunch to resume; stops when complete, retrigger for a new pass). Optional datetime overrides the cutoff and runs as a one-off (bypasses resume tracking). Optional blogname restricts the run to that blog (exact, case-sensitive match) and also runs as a one-off; e.g. \"--collect 1 zomb-eh\". datetime and blogname may be given in either order - use --blog=name if a blog name would otherwise parse as a date. Add --force to ignore the periodic re-collect cooldown and re-queue already-collected posts immediately (mode 1 only).");
Console.WriteLine("--blogsR\t For each Note that is a REPLY, write blogname to file ");
@@ -460,7 +471,7 @@ namespace URLNotesGrabberCORE
Console.WriteLine("--likes\t Fetch likes: initial backfill for new blogs, incremental refresh for blogs past cooldown. Optional blog name forces single-blog run.");
Console.WriteLine("--force\t (with --likes) Ignore cooldown and refresh every fully-backfilled blog");
Console.WriteLine("--force\t Ignore refresh cooldowns: with --likes, refresh every fully-backfilled blog; with --collect 1, re-queue already-collected posts without waiting out their cooldown");
Console.WriteLine("--urldump\t Scan all posts' text columns and extract suspected URLs to configured file");
@@ -1047,6 +1058,15 @@ namespace URLNotesGrabberCORE
string reblogKey = post.reblog_key?.ToString() ?? ".";
string link = ".";
// No file backs a liked post, so the filename trick used everywhere
// else cannot apply here. GrabLikes requests npf=false, and in the
// legacy format `type` is the discriminator that decides which content
// fields a post carries -- singular there, mapped to our plural names.
// liked_posts is List<dynamic>, so this is resolved at runtime and a
// missing field yields null rather than a compile error; an absent or
// unrecognized value leaves the type NULL instead of guessing.
string? apiPostType = PostTypes.FromApiType(post.type?.ToString() as string);
// Only insert if any of the data contains strings from ContainsList
bool shouldInsert = false;
string matchedFieldName = string.Empty;
@@ -1090,7 +1110,8 @@ if (shouldInsert)
DataAccess.AddPost(authorBlog, postID, reblogURL, date, postURL, slug, reblogKey,
reblogName, summary, quote, body, tags, link, photoURL,
photoCaption, downloadedFiles, audioCaption, question, answer,
title, hasImage, true, rootBlogName: rootBlogName, rootURL: rootURL);
title, hasImage, true, rootBlogName: rootBlogName, rootURL: rootURL,
postType: apiPostType);
}
}
@@ -1329,15 +1350,17 @@ if (shouldInsert)
// blipping on one post. Past this, skipping post-by-post would just hammer a closed door.
const int MaxConsecutiveTransient = 10;
static async Task<int> CollectNotes(string outPath, bool withoutNotesOnly = true, DateTime? beforeDate = null, bool managedRun = false, string? blogName = null)
static async Task<int> CollectNotes(string outPath, bool withoutNotesOnly = true, DateTime? beforeDate = null, bool managedRun = false, string? blogName = null, bool ignoreRefreshCooldown = false)
{
List<Tuple<string, long, long, long>> posts = DataAccess.GetPosts(withoutNotesOnly, beforeDate, blogName);
List<Tuple<string, long, long, long>> posts = DataAccess.GetPosts(withoutNotesOnly, beforeDate, blogName, ignoreRefreshCooldown);
if (posts.Count == 0 && !string.IsNullOrWhiteSpace(blogName))
{
// BlogName is matched exactly, so a typo or a case mismatch looks identical to "nothing
// left to collect". Say so rather than reporting a silent, instant success.
Console.WriteLine($"No posts to collect for blog '{blogName}'. Either it is fully collected, or the name does not match a stored blog (the match is case-sensitive).");
if (withoutNotesOnly && !ignoreRefreshCooldown)
Console.WriteLine("Already-collected posts are re-queued only once their cooldown elapses; add --force to re-collect them now.");
return 0;
}
@@ -1435,7 +1458,7 @@ if (shouldInsert)
}
// Re-fetch the updated list after processing the current post
posts = DataAccess.GetPosts(withoutNotesOnly, beforeDate, blogName);
posts = DataAccess.GetPosts(withoutNotesOnly, beforeDate, blogName, ignoreRefreshCooldown);
}
}
@@ -1508,7 +1531,15 @@ if (shouldInsert)
string normalizedDirectoryName = NormalizeBlogFolderName(new DirectoryInfo(path).Name);
bool isAtOrAfterStart = string.IsNullOrWhiteSpace(startFromBlogName) || string.Compare(normalizedDirectoryName, startFromBlogName, StringComparison.OrdinalIgnoreCase) >= 0;
// The filename is the post type ("texts.txt" -> "texts"), so only the eight
// known export files are post sources. Everything else in a blog folder is
// either not a post file at all (README.txt, url lists, triage scratch) or
// is our own derived output -- Unknown.txt above all, which must never be
// read back in as a source or it perpetuates itself.
string? filePostType = PostTypes.FromFileName(file);
if (file.EndsWith(".txt", StringComparison.OrdinalIgnoreCase)
&& filePostType != null
&& (string.IsNullOrEmpty(blogName) || path.IndexOf(blogName, StringComparison.OrdinalIgnoreCase) >= 0)
&& isAtOrAfterStart)
{
@@ -1541,7 +1572,7 @@ if (shouldInsert)
DataAccess.AddPost(curDir, long.Parse(reblog.postID), reblog.reblogURL, reblog.date, reblog.postURL, reblog.slug, reblog.reblogKey,
reblog.reblogName, reblog.summary, reblog.quote, reblog.body, reblog.tags, reblog.link, reblog.photoURL,
reblog.photoCaption, reblog.downloadedFiles, reblog.audioCaption, reblog.question, reblog.answer,
reblog.title, false, rootURL: reblog.rootURL);
reblog.title, false, rootURL: reblog.rootURL, postType: filePostType);
recordImportStopwatch.Stop();
postsAdded++;
@@ -1689,7 +1720,7 @@ if (shouldInsert)
DataAccess.AddPost(curDir, long.Parse(reblog.postID), reblog.reblogURL, reblog.date, reblog.postURL, reblog.slug, reblog.reblogKey,
reblog.reblogName, reblog.summary, reblog.quote, reblog.body, reblog.tags, reblog.link, reblog.photoURL,
reblog.photoCaption, reblog.downloadedFiles, reblog.audioCaption, reblog.question, reblog.answer,
reblog.title, true, rootURL: reblog.rootURL);
reblog.title, true, rootURL: reblog.rootURL, postType: filePostType);
recordImportStopwatch.Stop();
postsAdded++;