Author SHA1 Message Date
jim f541ec4260 fix: correctly detect rate-limit state for single-key API pools
IsAllRateLimited() short-circuited true for any pool with 0 or 1
keys, with minRetrySeconds left at 0 regardless of whether that key
was actually rate-limited. SleepUntilAnyAvailable() checks
"minRetry <= 0" to decide whether to skip sleeping, so with exactly
one key it always skipped the wait and let callers hammer the API
again immediately after a 429, even mid-cooldown.

The per-key loop already computes this correctly for any key count;
the special case only needs to cover the true no-keys edge case,
where there's nothing to wait on.
2026-06-30 21:22:12 -05:00
jim f549f020e1 refactor: standardize SQLiteConnection disposal via using; guard config
Replace the try/finally { connection.Close(); } pattern used across
most of DataAccess.cs with using declarations, so disposal happens
automatically and can't be skipped by a future edit that adds an
early return before the finally. Left the shared-connection
(ownsConnection) call sites alone since those intentionally outlive
a single method call.

Also drop a stray unused `using static ... JSType` import, and make
a missing ContainsList config setting fail with a clear
InvalidOperationException instead of a NullReferenceException from
Split(',') on null.
2026-06-30 20:54:18 -05:00
jim 0ff80a0fd3 fix: parameterize AddPost fallback UPDATE, guard args indexing
Posts.hasImage/DateModified fallback update built its WHERE clause via
raw string concatenation of blogName/postID, unlike every other query
in this method — a blog name containing a single quote would break or
inject into the query. Switch it to parameters.

--parse, --blogsO, and --bop indexed args[1..3] before checking
args.Length, so a missing argument threw IndexOutOfRangeException
instead of hitting the intended usage message.
2026-06-30 20:41:51 -05:00
jimandClaude Opus 4.8 4df73367fb BREAKING: switch all multi-char commands to POSIX --double-dash
Rename every multi-character option/command from single-dash to double-dash (--likes, --collect, --force, etc.) to follow the POSIX long-option convention. Single-character short options (-h, -V, -?) keep their single dash, as POSIX prescribes.

Breaking: existing invocations/scripts using single-dash forms now report Unknown Command and must be updated. Run profile (launchSettings.json) and CLI docs (copilot-instructions.md) updated to match.

Co-Authored-By: Claude Opus 4.8 <[email protected]>
2026-06-09 15:26:38 -05:00
jimandClaude Opus 4.8 a437fa87d3 Document -post and -bop commands in --help
These two commands were handled by the switch but never listed in help.
--help now covers every command the program accepts.

Co-Authored-By: Claude Opus 4.8 <[email protected]>
2026-06-09 15:19:51 -05:00
jimandClaude Opus 4.8 32a1583efd Document exit-status codes in --help output
The new 0/1/2 exit codes had no footprint in --help; add an Exit status
line so the documented behavior matches what the program now returns.

Co-Authored-By: Claude Opus 4.8 <[email protected]>
2026-06-09 15:18:25 -05:00
jimandClaude Opus 4.8 03676432bd Add POSIX-friendly CLI handling: --, --help/--version, exit codes
Keep the existing single-dash switch style and case-insensitive matching,
but add the cheap, non-breaking POSIX wins:

- `--` end-of-options: tokens after a bare `--` are treated as literal operands
- `--help`/`-h` (alongside `-?`) and `-V`/`--version`
- Main returns a real exit code: 2 for usage errors, propagates handler
  return codes, and a top-level catch yields a quiet 1 on unhandled errors

Co-Authored-By: Claude Opus 4.8 <[email protected]>
2026-06-09 15:12:10 -05:00
jim 8d6b9212c1 Remove binaries, batch script; add DB schema verifier
Removed outdated binary files and the `run_500_times.bat` script, which automated repetitive runs of `URLNotesGrabberCORE`. The batch script is no longer needed or has been replaced.

Added `verify-db-schema.sql`, a new script to verify and align the SQLite database schema (`TL.db`) with the application's expected schema. The script includes:
- A verification section to identify missing/extra columns or tables.
- An optional fix section with `ALTER TABLE` statements to add missing columns.

The SQL script ensures database compatibility while preserving data integrity and avoiding destructive operations.
2026-06-03 15:54:22 -05:00
jimandClaude Opus 4.8 18f172fe96 Make -collect 0 a resumable, single-pass full re-check
Mode 0 (full re-check) previously reset its cutoff to now on every
launch, so an interrupted run restarted from scratch, and a post that
kept returning a non-Success/non-NotFound status could loop forever.

- Add single-row CollectRunState table + accessors (EnsureCollectRunStateTableExists,
  GetCollectRunState, BeginCollectRun, CompleteCollectRun) mirroring the
  ApiKeyPoolMeta pattern, to persist a frozen run cutoff and completion flag.
- -collect 0 with no explicit date is now a managed run: resume against the
  stored cutoff if a run is in progress, else start a new run; mark complete
  when the pass finishes so the next launch starts fresh. Explicit-date and
  mode 1 behavior unchanged.
- CollectNotes makes a single attempt pass via an in-process attempted set;
  FAILURE/UNKNOWN are logged once, TooManyRequests/no-lease aborts without
  completing so a later launch resumes.

Co-Authored-By: Claude Opus 4.8 <[email protected]>
2026-06-03 15:48:26 -05:00
jimandClaude Opus 4.8 b576a9cdf3 fix: make -revert scan the PathInput tree like the no-parameter run
-revert was DB-driven, searching each blog's Blogs.TTFolderPath
non-recursively for *.bak. That tree differs from the no-parameter run,
which recursively walks PathInput. Rewrite RevertMode to recursively walk
PathInput (filesystem-only, no DB), with the optional [blogname] argument
now filtering by path substring. Restore mechanics unchanged.

Co-Authored-By: Claude Opus 4.8 <[email protected]>
2026-05-28 16:30:26 -05:00
jimandClaude Opus 4.8 5973920894 feat: add -revert mode to restore *.bak back to *.txt
Inverse of -output. For each blog with a TTFolderPath, restores every
*.bak over its *.txt, first preserving the current *.txt as the
next-free *.bkN, then consuming the *.bak. Confirms before running and
supports an optional single-blog filter.

Co-Authored-By: Claude Opus 4.8 <[email protected]>
2026-05-28 16:11:35 -05:00
jim 494d6aa2d4 Shortened sleep 2026-05-19 16:48:44 -05:00
jim 18c5ac5401 fix: correct foreach syntax for .NET 8 compatibility 2026-05-19 16:27:38 -05:00
jim 21e848efb7 feat: add [X remaining] progress counter to likes mode output 2026-05-19 15:48:18 -05:00
jim 24c5449e0c No longer copies db to output directory 2026-05-18 21:32:01 -05:00
jimandClaude Opus 4.7 c55569eadf chore: log folder transitions during -ingest
The every-50-file progress line wasn't enough to know which blog was
currently being processed on a long run. Now -ingest prints a
"entering folder: <name>" line whenever the source folder changes, and
the every-50 progress line also prefixes the folder name.

Co-Authored-By: Claude Opus 4.7 <[email protected]>
2026-05-18 15:40:27 -05:00
jimandClaude Opus 4.7 cf3f97ddc4 feat: add single-blog filter to -ingest
`-ingest <blogname>` now restricts the run to one blog's folder, mirroring
the existing -parse <blogname> ergonomics. `-ingest` with no arg still
processes every blog under appSettings:PathTTRoot (or PathInput fallback).

Breaking vs 3aff849: the first positional arg is interpreted as a blog
name, not a path. Configure the root via appSettings:PathTTRoot.

Co-Authored-By: Claude Opus 4.7 <[email protected]>
2026-05-18 15:36:42 -05:00
jimandClaude Opus 4.7 3aff849216 feat: merge ThreeTxtFileHelper into URLNotesGrabberCORE
Folds the standalone ThreeTxtFileHelper tool into URLNotesGrabberCORE so
text-file ingest/output/correct lives alongside the API scraper. Adds
new flags -ingest, -output, -correct (with -apply), -updatepaths, and a
one-time -importposts <posts.db> migration.

Schema: Blogs.TTFolderPath and Posts.PostType are added by an idempotent
migration. On (BlogName, PostID) collisions, content columns are
overwritten while engagement columns (ByLikes, RootBlogName, RootURL,
HasNotesGathered, NotFound, NotesGatheredDateTime, Likes*) are preserved.

Co-Authored-By: Claude Opus 4.7 <[email protected]>
2026-05-18 12:43:17 -05:00
jim 5781e121d2 Merge branch 'claude/gifted-dhawan-7e5fc7' 2026-05-16 14:02:07 -05:00
jimandClaude Sonnet 4.6 e27190e4f9 add incremental refresh + cooldown to -likes mode
After backfill completes for a blog, -likes can now pick up only newer
likes instead of being a one-shot pull. Tracks a per-blog
liked_timestamp high-water mark and stops the refresh walk once it
crosses the stored mark. A configurable cooldown (LikesRefreshCooldownDays,
default 7) gates which blogs are re-checked on each run. -force bypasses
the cooldown.

The migration adds three columns to Blogs (LikesNewestTimestamp,
LikesLastRefreshed, LikesLastNewCount) and does a one-time reset of all
likes tracking state so the new high-water mark starts from a clean
baseline. Existing ByLikes posts remain; the UNIQUE constraint absorbs
re-inserts during the first re-backfill.

Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
2026-05-16 14:00:45 -05:00
jim 7c667bd579 Merge branch 'claude/eager-hugle-2f3520' 2026-05-15 11:53:34 -05:00
jimandClaude Opus 4.7 2634ff8967 speed up no-args directory import with shared connection + lazy blog cache
AddBlog/AddPost/UpdatePost/UpdatePostSetDate now reuse a single SQLiteConnection
when an import session is active, instead of opening and closing one per call.
AddBlog also short-circuits on an in-memory HashSet of blog names already
attempted this run. Other entry points are unaffected since they never call
BeginImportSession.

Co-Authored-By: Claude Opus 4.7 <[email protected]>
2026-05-15 11:51:57 -05:00
20 changed files with 2449 additions and 349 deletions
+9 -9
View File
@@ -22,18 +22,18 @@ This document provides essential context for AI agents working with URLNotesGrab
```powershell ```powershell
dotnet build dotnet build
dotnet run # Process all files in input directory dotnet run # Process all files in input directory
dotnet run -- -parse [blogname] # Process specific blog dotnet run -- --parse [blogname] # Process specific blog
dotnet run -- -test [blogname] [postID] # Test API for specific post dotnet run -- --test [blogname] [postID] # Test API for specific post
``` ```
### Command-Line Interface ### Command-Line Interface
- `-parse [blogname]`: Parse text files for specific blog - `--parse [blogname]`: Parse text files for specific blog
- `-test [blogname] [postID]`: Test API note collection - `--test [blogname] [postID]`: Test API note collection
- `-posts`: Export post blogs to file - `--posts`: Export post blogs to file
- `-blogs`: Export blog list to file - `--blogs`: Export blog list to file
- `-collect`: Collect notes for all posts in DB - `--collect`: Collect notes for all posts in DB
- `-blogsR`: Export reply blogs to file - `--blogsR`: Export reply blogs to file
- `-blogsO [start] [stop]`: Export blogs within range - `--blogsO [start] [stop]`: Export blogs within range
## Project Conventions ## Project Conventions
BIN
View File
Binary file not shown.
+24 -5
View File
@@ -1,4 +1,4 @@
<?xml version="1.0" encoding="UTF-8"?><sqlb_project><db path="C:/Users/jim/Nextcloud/C#/URLNotesGrabberCORE/URLNotesGrabberCORE/TL.db" readonly="0" foreign_keys="1" case_sensitive_like="0" temp_store="0" wal_autocheckpoint="1000" synchronous="2"/><attached/><window><main_tabs open="structure browser pragmas query" current="3"/></window><tab_structure><column_width id="0" width="300"/><column_width id="1" width="0"/><column_width id="2" width="100"/><column_width id="3" width="4253"/><column_width id="4" width="0"/><expanded_item id="0" parent="1"/><expanded_item id="1" parent="1"/><expanded_item id="2" parent="1"/><expanded_item id="3" parent="1"/></tab_structure><tab_browse><table title="ApiKeyPoolMeta" custom_title="0" dock_id="4" table="4,14:mainApiKeyPoolMeta"/><dock_state state="000000ff00000000fd00000001000000020000077400000387fc0100000006fb000000160064006f0063006b00420072006f00770073006500310100000000000004a10000000000000000fb000000160064006f0063006b00420072006f00770073006500320100000000000004a10000000000000000fb000000160064006f0063006b00420072006f00770073006500330100000000000004a10000000000000000fb000000160064006f0063006b00420072006f00770073006500350100000000000005f40000000000000000fc00000000000007740000000000fffffffaffffffff0100000001fb000000160064006f0063006b00420072006f00770073006500340000000000ffffffff0000000000000000fb000000160064006f0063006b00420072006f00770073006500340100000000ffffffff0000011e00ffffff000007740000000000000004000000040000000800000008fc00000000"/><default_encoding codec=""/><browse_table_settings/></tab_browse><tab_sql><sql name="SQL 1">UPDATE Posts <?xml version="1.0" encoding="UTF-8"?><sqlb_project><db path="C:/Users/jim/Nextcloud/C#/URLNotesGrabberCORE/URLNotesGrabberCORE/TL.db" readonly="0" foreign_keys="1" case_sensitive_like="0" temp_store="0" wal_autocheckpoint="1000" synchronous="2"/><attached/><window><main_tabs open="structure browser pragmas query" current="3"/></window><tab_structure><column_width id="0" width="300"/><column_width id="1" width="0"/><column_width id="2" width="100"/><column_width id="3" width="4305"/><column_width id="4" width="0"/><expanded_item id="0" parent="1"/><expanded_item id="1" parent="1"/><expanded_item id="2" parent="1"/><expanded_item id="3" parent="1"/></tab_structure><tab_browse><table title="Posts" custom_title="0" dock_id="4" table="4,5:mainPosts"/><dock_state state="000000ff00000000fd00000001000000020000077200000379fc0100000006fb000000160064006f0063006b00420072006f00770073006500310100000000000004a10000000000000000fb000000160064006f0063006b00420072006f00770073006500320100000000000004a10000000000000000fb000000160064006f0063006b00420072006f00770073006500330100000000000004a10000000000000000fb000000160064006f0063006b00420072006f00770073006500350100000000000005f40000000000000000fb000000160064006f0063006b00420072006f00770073006500340100000000000007720000011700fffffffb000000160064006f0063006b00420072006f00770073006500340100000000000005f40000000000000000000007720000000000000004000000040000000800000008fc00000000"/><default_encoding codec=""/><browse_table_settings><table schema="main" name="ApiKeyPoolMeta" show_row_id="0" encoding="" plot_x_axis="" unlock_view_pk="_rowid_" freeze_columns="0"><sort/><column_widths><column index="1" value="29"/><column index="2" value="64"/></column_widths><filter_values/><conditional_formats/><row_id_formats/><display_formats/><hidden_columns/><plot_y_axes/><global_filter/></table><table schema="main" name="Blogs" show_row_id="0" encoding="" plot_x_axis="" unlock_view_pk="_rowid_" freeze_columns="0"><sort/><column_widths><column index="1" value="257"/><column index="2" value="95"/><column index="3" value="54"/><column index="4" value="156"/><column index="5" value="51"/><column index="6" value="71"/><column index="7" value="85"/><column index="8" value="156"/><column index="9" value="156"/></column_widths><filter_values/><conditional_formats/><row_id_formats/><display_formats/><hidden_columns/><plot_y_axes/><global_filter/></table><table schema="main" name="Posts" show_row_id="0" encoding="" plot_x_axis="" unlock_view_pk="_rowid_" freeze_columns="0"><sort><column index="28" mode="1"/></sort><column_widths><column index="1" value="241"/><column index="2" value="148"/><column index="3" value="126"/><column index="4" value="300"/><column index="5" value="75"/><column index="6" value="187"/><column index="7" value="159"/><column index="8" value="75"/><column index="9" value="300"/><column index="10" value="300"/><column index="11" value="78"/><column index="12" value="249"/><column index="13" value="300"/><column index="14" value="53"/><column index="15" value="300"/><column index="16" value="300"/><column index="17" value="41"/><column index="18" value="75"/><column index="19" value="96"/><column index="20" value="300"/><column index="21" value="96"/><column index="22" value="300"/><column index="23" value="300"/><column index="24" value="42"/><column index="25" value="60"/><column index="26" value="218"/><column index="27" value="920"/><column index="28" value="156"/><column index="29" value="156"/></column_widths><filter_values><column index="24" value="=1"/><column index="28" value="&gt;2026-05-06 20:00:01"/></filter_values><conditional_formats/><row_id_formats/><display_formats/><hidden_columns/><plot_y_axes/><global_filter/></table></browse_table_settings></tab_browse><tab_sql><sql name="SQL 1">UPDATE Posts
SET HasNotesGathered = 0 SET HasNotesGathered = 0
WHERE (BlogName, PostID) IN ( WHERE (BlogName, PostID) IN (
SELECT p.BlogName, p.PostID SELECT p.BlogName, p.PostID
@@ -31,9 +31,10 @@ blogname in
)</sql><sql name="New Notes">select RootBlogName, PostID, NoteBlogName || '.tumblr.com' as NoteBlogName, DatetimeCrawled, TimeStamp, type, RootBlogName || '.tumblr.com/post/' || postid, datetime(timestamp, 'unixepoch') )</sql><sql name="New Notes">select RootBlogName, PostID, NoteBlogName || '.tumblr.com' as NoteBlogName, DatetimeCrawled, TimeStamp, type, RootBlogName || '.tumblr.com/post/' || postid, datetime(timestamp, 'unixepoch')
from Notes from Notes
where --type like 'r%' and where
DatetimeCrawled &lt;&gt; '2026-04-30 09:25:43' DatetimeCrawled &gt; '2026-05-14 02:50:05' --and type like 'r%'
order by DatetimeCrawled desc, TimeStamp desc</sql><sql name="Pull Blogs">SELECT distinct order by DatetimeCrawled desc</sql><sql name="Pull Blogs*">SELECT distinct
'''' || blogname || ''',',
blogs.* blogs.*
, blogname || '.tumblr.com' , blogname || '.tumblr.com'
FROM FROM
@@ -45,4 +46,22 @@ WHERE
order by order by
Notes.Type desc, Notes.Type desc,
DateAdded desc DateAdded desc
LIMIT 100;</sql><current_tab id="1"/></tab_sql></sqlb_project> LIMIT 100;</sql><sql name="SQL 7">WITH ReplyCounts AS (
SELECT
NoteBlogName,
COUNT(DISTINCT replyText) AS DistinctReplyCount
FROM Notes
where replyText &lt;&gt; '.'
GROUP BY NoteBlogName
)
SELECT
n.RootBlogName || '.tumblr.com/post/' || n.PostID AS PostURL, postid,
n.NoteBlogName,
n.replyText,
c.DistinctReplyCount
FROM Notes n
JOIN ReplyCounts c ON n.NoteBlogName = c.NoteBlogName
where replyText &lt;&gt; '.' and type &lt;&gt; 'reply'
--AND N.NoteBlogName NOT IN ( 'roadblocker21', 'thesaddemon666', 'edwardabbeyhoffman', 'tattedsoldier20', 'zomb-eh', 'animalistic13', 'indken', 'maccloud1592',
--'moss-wizard', 'supertrucker12682', 'exploringthrupics', 'padeyepete' )
order by c.DistinctReplyCount desc, n.NoteBlogName, n.DateModified desc, replyText, RootBlogName, PostID</sql><current_tab id="3"/></tab_sql></sqlb_project>
BIN
View File
Binary file not shown.
BIN
View File
Binary file not shown.
BIN
View File
Binary file not shown.
+277
View File
@@ -0,0 +1,277 @@
using Microsoft.Extensions.Configuration;
namespace URLNotesGrabberCORE
{
// Port of ThreeTxtFileHelper RunCorrectionMode (dry-run) and RunFullCorrectionMode (apply).
// Scans a BAK directory of .txt files, parses posts with multi-line field support,
// and either reports or applies content-column corrections to TL.db.Posts.
// Apply path only writes non-empty values (mirrors original ThreeTxtFileHelper semantics)
// and never touches engagement columns.
public static class CorrectMode
{
public static int Run(IConfiguration config, string[] args, bool applyChanges)
{
DataAccess.EnsureTTFileHelperColumnsExist();
// Resolve BAK path: explicit arg > appSettings:PathTTBackup > derive from PathTTRoot/PathInput
string? bakRootPath = args.Length > 0 ? args[0] : config["appSettings:PathTTBackup"];
if (string.IsNullOrWhiteSpace(bakRootPath))
{
string? root = config["appSettings:PathTTRoot"];
if (string.IsNullOrWhiteSpace(root)) root = config["appSettings:PathInput"];
if (!string.IsNullOrWhiteSpace(root))
bakRootPath = root.TrimEnd('\\', '/') + "_BAK\\";
}
if (string.IsNullOrWhiteSpace(bakRootPath) || !Directory.Exists(bakRootPath))
{
Console.WriteLine($"BAK directory not found: {bakRootPath}");
return 1;
}
Console.WriteLine($"BAK source path: {bakRootPath}");
Console.WriteLine(applyChanges
? "Correction mode: APPLY - non-empty fields from BAK overwrite DB columns\n"
: "Correction mode: Dry-run - reports multi-line field updates available\n");
string prefixesPath = config["appSettings:PathPrefixes"] ?? "prefixes.txt";
if (!Path.IsPathRooted(prefixesPath))
prefixesPath = Path.Combine(AppContext.BaseDirectory, prefixesPath);
if (!File.Exists(prefixesPath))
{
Console.WriteLine($"Prefixes file not found: {prefixesPath}");
return 1;
}
var allowedPrefixes = new HashSet<string>(File.ReadLines(prefixesPath), StringComparer.OrdinalIgnoreCase);
if (applyChanges)
{
Console.Write("WARNING: This will overwrite non-empty fields in matching posts from BAK files. Continue? (yes/no): ");
string? response = Console.ReadLine();
if (string.IsNullOrWhiteSpace(response) || !response.Equals("yes", StringComparison.OrdinalIgnoreCase))
{
Console.WriteLine("Operation cancelled.");
return 0;
}
}
var bakTxtFiles = new List<string>();
try
{
foreach (var dir in Directory.GetDirectories(bakRootPath, "*", SearchOption.AllDirectories))
bakTxtFiles.AddRange(Directory.GetFiles(dir, "*.txt"));
bakTxtFiles.AddRange(Directory.GetFiles(bakRootPath, "*.txt"));
}
catch (Exception ex)
{
Console.WriteLine($"Error scanning BAK directory: {ex.Message}");
return 1;
}
Console.WriteLine($"Found {bakTxtFiles.Count} file(s) in BAK directory\n");
int totalPostsFound = 0;
int postsWithUpdates = 0;
int postsUpdated = 0;
int postsNotFound = 0;
var correctionLog = new List<string>();
var updateLog = new List<string>();
foreach (string bakFile in bakTxtFiles)
{
Console.WriteLine($"Processing BAK file: {Path.GetFileName(bakFile)}");
try
{
var bakPosts = ParsePostsFromFile(bakFile, allowedPrefixes);
Console.WriteLine($" Found {bakPosts.Count} post(s) in this file");
foreach (var (postId, bakData) in bakPosts)
{
totalPostsFound++;
var dbPost = DataAccess.GetPostByIdAnyBlog(postId);
if (dbPost == null)
{
postsNotFound++;
continue;
}
if (applyChanges)
{
bool updated = DataAccess.UpdatePostContentFields(dbPost.BlogName, dbPost.PostId, bakData);
if (updated)
{
postsUpdated++;
updateLog.Add($"Post ID: {postId} - Updated from {Path.GetFileName(bakFile)}");
}
}
else
{
var updateList = BuildDryRunDiff(bakData, dbPost);
if (updateList.Count > 0)
{
postsWithUpdates++;
correctionLog.Add($"\nPost ID: {postId}");
correctionLog.Add($" File: {Path.GetFileName(bakFile)}");
correctionLog.Add($" Fields to update:");
correctionLog.AddRange(updateList);
}
}
}
}
catch (Exception ex)
{
Console.WriteLine($" ERROR processing file: {ex.Message}");
}
}
if (applyChanges)
{
Console.WriteLine($"\n========== CORRECTION COMPLETE ==========");
Console.WriteLine($"Total posts found in BAK files: {totalPostsFound}");
Console.WriteLine($"Posts updated in database: {postsUpdated}");
Console.WriteLine($"Posts not found in database: {postsNotFound}");
string logPath = config["appSettings:PathCorrectionApplied"] ?? "correction_applied.txt";
try
{
var logLines = new List<string>
{
$"Correction Applied: {DateTime.Now:yyyy-MM-dd HH:mm:ss}",
$"Total posts found in BAK files: {totalPostsFound}",
$"Posts updated in database: {postsUpdated}",
$"Posts not found in database: {postsNotFound}",
"",
"Updated Posts:"
};
logLines.AddRange(updateLog);
File.WriteAllLines(logPath, logLines);
Console.WriteLine($"Update log saved to: {logPath}");
}
catch (Exception ex)
{
Console.WriteLine($"Error writing log file: {ex.Message}");
}
}
else
{
Console.WriteLine($"\n========== CORRECTION REPORT (DRY RUN) ==========");
Console.WriteLine($"Total posts found in BAK files: {totalPostsFound}");
Console.WriteLine($"Posts with multi-line field updates available: {postsWithUpdates}");
if (correctionLog.Count > 0)
{
string logPath = config["appSettings:PathCorrectionReport"] ?? "correction_report.txt";
try
{
File.WriteAllLines(logPath, correctionLog);
Console.WriteLine($"\nDetailed report saved to: {logPath}");
}
catch (Exception ex)
{
Console.WriteLine($"Error writing report file: {ex.Message}");
}
}
else
{
Console.WriteLine("\nNo multi-line field updates found.");
}
Console.WriteLine("\nDry-run complete. No database changes were made.");
Console.WriteLine("If updates look correct, re-run with `-correct -apply` to apply changes.");
}
return 0;
}
private static List<string> BuildDryRunDiff(Dictionary<string, string> bakData, TTPostRecord dbPost)
{
var updates = new List<string>();
foreach (var (fieldName, bakValue) in bakData)
{
if (string.IsNullOrWhiteSpace(bakValue)) continue;
string? currentValue = fieldName.ToLowerInvariant() switch
{
"reblog url" => dbPost.ReblogUrl,
"date" => dbPost.Date,
"has image" => dbPost.HasImage,
"post url" => dbPost.PostUrl,
"slug" => dbPost.Slug,
"reblog key" => dbPost.ReblogKey,
"reblog name" => dbPost.ReblogName,
"summary" => dbPost.Summary,
"quote" => dbPost.Quote,
"body" => dbPost.Body,
"tags" => dbPost.Tags,
"link" => dbPost.Link,
"photo url" => dbPost.PhotoUrl,
"photo caption" => dbPost.PhotoCaption,
"downloaded files" => dbPost.DownloadedFiles,
"audio caption" => dbPost.AudioCaption,
"question" => dbPost.Question,
"answer" => dbPost.Answer,
"title" => dbPost.Title,
_ => null
};
if (bakValue != currentValue && bakValue.Contains('\n'))
updates.Add($" {fieldName}: [MULTILINE]");
}
return updates;
}
private static Dictionary<string, Dictionary<string, string>> ParsePostsFromFile(string filePath, HashSet<string> allowedPrefixes)
{
var posts = new Dictionary<string, Dictionary<string, string>>(StringComparer.OrdinalIgnoreCase);
var lines = File.ReadAllLines(filePath);
int lineIndex = 0;
string currentPostId = "";
var currentPostData = new Dictionary<string, string>(StringComparer.OrdinalIgnoreCase);
while (lineIndex < lines.Length)
{
string line = lines[lineIndex];
string searchText = line.Length > 25 ? line.Substring(0, 25) : line;
int colonIndex = searchText.IndexOf(": ");
if (colonIndex > 0)
{
string prefix = line.Substring(0, colonIndex).Trim();
if (!string.IsNullOrWhiteSpace(prefix) && allowedPrefixes.Contains(prefix))
{
if (string.Equals(prefix, "Post ID", StringComparison.OrdinalIgnoreCase))
{
if (!string.IsNullOrWhiteSpace(currentPostId) && currentPostData.Count > 0)
posts[currentPostId] = new Dictionary<string, string>(currentPostData, StringComparer.OrdinalIgnoreCase);
currentPostId = line.Substring(colonIndex + 2).Trim();
currentPostData.Clear();
lineIndex++;
continue;
}
var valueLines = new List<string> { line.Substring(colonIndex + 2).Trim() };
int nextLineIndex = lineIndex + 1;
while (nextLineIndex < lines.Length)
{
string nextLine = lines[nextLineIndex];
string nextSearch = nextLine.Length > 25 ? nextLine.Substring(0, 25) : nextLine;
int nextColon = nextSearch.IndexOf(": ");
if (nextColon > 0)
{
string nextPrefix = nextLine.Substring(0, nextColon).Trim();
if (!string.IsNullOrWhiteSpace(nextPrefix) && allowedPrefixes.Contains(nextPrefix))
break;
}
valueLines.Add(nextLine);
nextLineIndex++;
}
currentPostData[prefix] = string.Join("\n", valueLines);
lineIndex = nextLineIndex;
continue;
}
}
lineIndex++;
}
if (!string.IsNullOrWhiteSpace(currentPostId) && currentPostData.Count > 0)
posts[currentPostId] = new Dictionary<string, string>(currentPostData, StringComparer.OrdinalIgnoreCase);
return posts;
}
}
}
File diff suppressed because it is too large Load Diff
+219
View File
@@ -0,0 +1,219 @@
using System.Text.RegularExpressions;
using Microsoft.Extensions.Configuration;
namespace URLNotesGrabberCORE
{
// Port of ThreeTxtFileHelper RunIngestMode. Scans a root folder for .txt files,
// parses Tumblr-export fields (multi-line aware, prefix-driven), upserts each
// post into TL.db.Posts via DataAccess.UpsertPostFromTextFile.
public static class IngestMode
{
public static int Run(IConfiguration config, string[] args)
{
string? targetBlog = args.Length > 0 ? args[0]?.Trim() : null;
if (string.IsNullOrWhiteSpace(targetBlog)) targetBlog = null;
string? rootPath = config["appSettings:PathTTRoot"];
if (string.IsNullOrWhiteSpace(rootPath))
rootPath = config["appSettings:PathInput"];
if (string.IsNullOrWhiteSpace(rootPath))
{
Console.WriteLine("Ingest: no root path configured. Set appSettings:PathTTRoot or appSettings:PathInput.");
return 1;
}
if (!Directory.Exists(rootPath))
{
Console.WriteLine($"Directory not found: {rootPath}");
return 1;
}
DataAccess.EnsureTTFileHelperColumnsExist();
string prefixesPath = config["appSettings:PathPrefixes"] ?? "prefixes.txt";
if (!Path.IsPathRooted(prefixesPath))
prefixesPath = Path.Combine(AppContext.BaseDirectory, prefixesPath);
if (!File.Exists(prefixesPath))
{
Console.WriteLine($"Prefixes file not found: {prefixesPath}");
return 1;
}
var allowedPrefixes = new HashSet<string>(File.ReadLines(prefixesPath), StringComparer.OrdinalIgnoreCase);
Console.WriteLine($"Loaded {allowedPrefixes.Count} prefixes from {prefixesPath}");
Console.WriteLine($"========== Ingest Settings ==========");
Console.WriteLine($"Root path: {rootPath}");
Console.WriteLine($"Blog filter: {(targetBlog == null ? "(all blogs)" : targetBlog)}");
Console.WriteLine($"=====================================");
var txtFiles = new List<string>();
try
{
var dirs = Directory.GetDirectories(rootPath, "*", SearchOption.AllDirectories);
Console.WriteLine($"Found {dirs.Length} directories under root.");
foreach (var dir in dirs)
{
try { txtFiles.AddRange(Directory.GetFiles(dir, "*.txt")); }
catch (Exception ex) { Console.WriteLine($" Skipping {dir}: {ex.Message}"); }
}
txtFiles.AddRange(Directory.GetFiles(rootPath, "*.txt"));
}
catch (Exception ex)
{
Console.WriteLine($"Error scanning root: {ex.Message}");
return 1;
}
Console.WriteLine($"Processing {txtFiles.Count} .txt file(s)...");
int filesProcessed = 0;
int filesSkipped = 0;
int postsTouched = 0;
string? lastBlogFolder = null;
try
{
DataAccess.EnableImportModePragmas();
DataAccess.BeginImportSession();
foreach (string file in txtFiles)
{
try
{
string rawBlogName = Path.GetFileName(Path.GetDirectoryName(file) ?? "unknown");
string blogName = Regex.Replace(rawBlogName, @"_\d+$", "");
string postType = Path.GetFileNameWithoutExtension(file);
if (targetBlog != null && !string.Equals(blogName, targetBlog, StringComparison.OrdinalIgnoreCase))
{
filesSkipped++;
continue;
}
filesProcessed++;
if (lastBlogFolder != rawBlogName)
{
Console.WriteLine($"[{filesProcessed}/{txtFiles.Count}] >> entering folder: {rawBlogName}");
lastBlogFolder = rawBlogName;
}
else if (filesProcessed % 50 == 0)
{
Console.WriteLine($"[{filesProcessed}/{txtFiles.Count}] {rawBlogName}/{Path.GetFileName(file)}");
}
string currentPostId = "";
var currentPostData = new Dictionary<string, string>(StringComparer.OrdinalIgnoreCase);
void Flush()
{
if (!string.IsNullOrWhiteSpace(currentPostId) && currentPostData.Count > 0)
{
UpsertPostFromParsedData(blogName, currentPostId, postType, currentPostData);
postsTouched++;
}
}
var lines = File.ReadAllLines(file);
int lineIndex = 0;
while (lineIndex < lines.Length)
{
string line = lines[lineIndex];
string searchText = line.Length > 25 ? line.Substring(0, 25) : line;
int colonIndex = searchText.IndexOf(": ");
if (colonIndex > 0)
{
string prefix = line.Substring(0, colonIndex).Trim();
if (!string.IsNullOrWhiteSpace(prefix) && allowedPrefixes.Contains(prefix))
{
if (string.Equals(prefix, "Post ID", StringComparison.OrdinalIgnoreCase))
{
Flush();
currentPostId = line.Substring(colonIndex + 2).Trim();
currentPostData.Clear();
lineIndex++;
continue;
}
var valueLines = new List<string> { line.Substring(colonIndex + 2).Trim() };
int nextLineIndex = lineIndex + 1;
while (nextLineIndex < lines.Length)
{
string nextLine = lines[nextLineIndex];
string nextSearch = nextLine.Length > 25 ? nextLine.Substring(0, 25) : nextLine;
int nextColon = nextSearch.IndexOf(": ");
if (nextColon > 0)
{
string nextPrefix = nextLine.Substring(0, nextColon).Trim();
if (!string.IsNullOrWhiteSpace(nextPrefix) && allowedPrefixes.Contains(nextPrefix))
break;
}
valueLines.Add(nextLine);
nextLineIndex++;
}
currentPostData[prefix] = string.Join("\n", valueLines);
lineIndex = nextLineIndex;
continue;
}
}
lineIndex++;
}
Flush();
}
catch (Exception ex)
{
Console.WriteLine($" ERROR processing file {file}: {ex.Message}");
}
}
}
finally
{
DataAccess.EndImportSession();
DataAccess.RestoreImportModePragmas();
}
Console.WriteLine($"\nIngest complete. Files processed: {filesProcessed}. Files skipped (blog filter): {filesSkipped}. Posts touched: {postsTouched}.");
return 0;
}
private static void UpsertPostFromParsedData(string blogName, string postId, string postType, Dictionary<string, string> data)
{
string? G(string key) => data.TryGetValue(key, out var v) ? v : null;
string? hasImageStr = G("Has Image");
bool hasImage = !string.IsNullOrWhiteSpace(hasImageStr)
&& (hasImageStr.Equals("true", StringComparison.OrdinalIgnoreCase)
|| hasImageStr == "1"
|| hasImageStr.Equals("yes", StringComparison.OrdinalIgnoreCase));
DataAccess.UpsertPostFromTextFile(
blogName: blogName,
postID: postId,
reblogURL: G("reblog URL"),
postDate: G("Date"),
postURL: G("Post URL"),
slug: G("Slug"),
reblogKey: G("Reblog Key"),
reblogName: G("Reblog Name"),
summary: G("Summary"),
quote: G("Quote"),
body: G("Body"),
tags: G("Tags"),
link: G("Link"),
photoURL: G("Photo URL"),
photoCaption: G("Photo Caption"),
downloadedFiles: G("Downloaded Files"),
audioCaption: G("Audio Caption"),
question: G("Question"),
answer: G("Answer"),
title: G("Title"),
postType: postType,
hasImage: hasImage);
}
}
}
@@ -0,0 +1,147 @@
using System.Data.SQLite;
namespace URLNotesGrabberCORE
{
// One-time migration: opens a legacy ThreeTxtFileHelper posts.db, copies its
// Blog + PostData rows into the merged TL.db via DataAccess.
// Conflict rule on (BlogName, PostId): ThreeTxtFileHelper wins on the 22 content
// columns + PostType + DateModified (handled inside UpsertPostFromTextFile).
// Engagement columns in TL.db (ByLikes, RootBlogName, RootURL, HasNotesGathered,
// NotFound, NotesGatheredDateTime) are preserved.
public static class LegacyPostsDbImporter
{
public static int Run(string legacyDbPath)
{
if (string.IsNullOrWhiteSpace(legacyDbPath))
{
Console.WriteLine("LegacyPostsDbImporter: path to legacy posts.db is required.");
return 1;
}
if (!File.Exists(legacyDbPath))
{
Console.WriteLine($"Legacy posts.db not found at: {legacyDbPath}");
return 1;
}
DataAccess.EnsureTTFileHelperColumnsExist();
Console.WriteLine($"Reading legacy posts.db: {legacyDbPath}");
int blogsCopied = 0;
int postsUpserted = 0;
int errors = 0;
try
{
using var src = new SQLiteConnection("Data Source=" + legacyDbPath + ";Read Only=True;");
src.Open();
// 1) Copy Blogs (BlogName + TTFolderPath)
using (var cmd = new SQLiteCommand("SELECT BlogName, TTFolderPath FROM Blogs", src))
using (var reader = cmd.ExecuteReader())
{
while (reader.Read())
{
string blogName = reader.IsDBNull(0) ? string.Empty : reader.GetString(0);
string? ttFolderPath = reader.IsDBNull(1) ? null : reader.GetString(1);
if (string.IsNullOrWhiteSpace(blogName)) continue;
try
{
DataAccess.SetBlogTTFolderPath(blogName, ttFolderPath);
blogsCopied++;
}
catch (Exception ex)
{
errors++;
Console.WriteLine($" Blog copy failed for '{blogName}': {ex.Message}");
}
}
}
Console.WriteLine($" Blogs copied: {blogsCopied}");
// 2) Copy Posts
try
{
DataAccess.EnableImportModePragmas();
DataAccess.BeginImportSession();
string sql = @"SELECT BlogName, PostId, ReblogUrl, Date, HasImage, PostUrl, Slug,
ReblogKey, ReblogName, Summary, Quote, Body, Tags, Link,
PhotoUrl, PhotoCaption, DownloadedFiles, AudioCaption,
Question, Answer, Title, PostType
FROM Posts";
using var cmd = new SQLiteCommand(sql, src);
using var reader = cmd.ExecuteReader();
while (reader.Read())
{
try
{
string blogName = reader.IsDBNull(0) ? string.Empty : reader.GetString(0);
string postId = reader.IsDBNull(1) ? string.Empty : reader.GetString(1);
if (string.IsNullOrWhiteSpace(blogName) || string.IsNullOrWhiteSpace(postId)) continue;
string? hasImageRaw = reader.IsDBNull(4) ? null : reader.GetValue(4)?.ToString();
bool hasImage = !string.IsNullOrWhiteSpace(hasImageRaw)
&& (hasImageRaw.Equals("true", StringComparison.OrdinalIgnoreCase)
|| hasImageRaw == "1"
|| hasImageRaw.Equals("yes", StringComparison.OrdinalIgnoreCase));
DataAccess.UpsertPostFromTextFile(
blogName: blogName,
postID: postId,
reblogURL: reader.IsDBNull(2) ? null : reader.GetString(2),
postDate: reader.IsDBNull(3) ? null : reader.GetString(3),
postURL: reader.IsDBNull(5) ? null : reader.GetString(5),
slug: reader.IsDBNull(6) ? null : reader.GetString(6),
reblogKey: reader.IsDBNull(7) ? null : reader.GetString(7),
reblogName: reader.IsDBNull(8) ? null : reader.GetString(8),
summary: reader.IsDBNull(9) ? null : reader.GetString(9),
quote: reader.IsDBNull(10) ? null : reader.GetString(10),
body: reader.IsDBNull(11) ? null : reader.GetString(11),
tags: reader.IsDBNull(12) ? null : reader.GetString(12),
link: reader.IsDBNull(13) ? null : reader.GetString(13),
photoURL: reader.IsDBNull(14) ? null : reader.GetString(14),
photoCaption: reader.IsDBNull(15) ? null : reader.GetString(15),
downloadedFiles: reader.IsDBNull(16) ? null : reader.GetString(16),
audioCaption: reader.IsDBNull(17) ? null : reader.GetString(17),
question: reader.IsDBNull(18) ? null : reader.GetString(18),
answer: reader.IsDBNull(19) ? null : reader.GetString(19),
title: reader.IsDBNull(20) ? null : reader.GetString(20),
postType: reader.IsDBNull(21) ? null : reader.GetString(21),
hasImage: hasImage);
postsUpserted++;
if (postsUpserted % 500 == 0)
Console.WriteLine($" ... {postsUpserted} posts upserted");
}
catch (Exception ex)
{
errors++;
if (errors < 20)
Console.WriteLine($" Post upsert error: {ex.Message}");
}
}
}
finally
{
DataAccess.EndImportSession();
DataAccess.RestoreImportModePragmas();
}
Console.WriteLine($" Posts upserted: {postsUpserted}");
}
catch (Exception ex)
{
Console.WriteLine($"Fatal error reading legacy posts.db: {ex.Message}");
return 1;
}
Console.WriteLine($"\n========== Legacy import summary ==========");
Console.WriteLine($"Blogs copied: {blogsCopied}");
Console.WriteLine($"Posts upserted: {postsUpserted}");
Console.WriteLine($"Errors: {errors}");
return errors == 0 ? 0 : 2;
}
}
}
+136
View File
@@ -0,0 +1,136 @@
using Microsoft.Extensions.Configuration;
namespace URLNotesGrabberCORE
{
// Port of ThreeTxtFileHelper RunOutputMode + WritePostToFile + RenameExistingTxtFilesToBak.
// For each Blog with a TTFolderPath, renames any existing .txt files in that folder to .bak,
// then writes one .txt per PostType containing all posts of that type (date-sorted, fixed
// field order). Reads from TL.db via DataAccess.GetAllPostsForBlog.
public static class OutputMode
{
public static int Run(IConfiguration config)
{
DataAccess.EnsureTTFileHelperColumnsExist();
var blogs = DataAccess.GetAllBlogsWithTTFolderPath();
Console.WriteLine($"Found {blogs.Count} blog(s) to process.");
foreach (var (blogName, ttFolderPath) in blogs)
{
Console.WriteLine($"\nProcessing blog: {blogName}");
if (string.IsNullOrWhiteSpace(ttFolderPath) || !Directory.Exists(ttFolderPath))
{
Console.WriteLine($" TTFolderPath does not exist or is not set. Skipping.");
continue;
}
Console.WriteLine($" TTFolderPath: {ttFolderPath}");
try
{
foreach (var bakFile in Directory.GetFiles(ttFolderPath, "*.bak"))
File.Delete(bakFile);
}
catch (Exception ex)
{
Console.WriteLine($" Error deleting .bak files: {ex.Message}");
}
RenameExistingTxtFilesToBak(ttFolderPath);
var posts = DataAccess.GetAllPostsForBlog(blogName);
Console.WriteLine($" Found {posts.Count} post(s) for this blog.");
var grouped = posts.GroupBy(p => p.PostType ?? "Unknown");
foreach (var typeGroup in grouped)
{
string postType = typeGroup.Key ?? "Unknown";
string outputFilePath = Path.Combine(ttFolderPath, $"{postType}.txt");
var ordered = typeGroup.OrderBy(p => p.Date).ToList();
Console.WriteLine($" Writing {ordered.Count} post(s) to {postType}.txt");
using var writer = new StreamWriter(outputFilePath, false, System.Text.Encoding.UTF8);
bool isFirst = true;
foreach (var post in ordered)
{
if (!isFirst)
{
writer.WriteLine();
writer.WriteLine();
}
WritePostToFile(writer, post);
isFirst = false;
}
}
}
Console.WriteLine("\nOutput mode complete.");
return 0;
}
private static void RenameExistingTxtFilesToBak(string folderPath)
{
try
{
foreach (var txtFile in Directory.GetFiles(folderPath, "*.txt"))
{
string bakPath = Path.ChangeExtension(txtFile, ".bak");
if (File.Exists(bakPath)) File.Delete(bakPath);
File.Move(txtFile, bakPath, overwrite: true);
}
}
catch (Exception ex)
{
Console.WriteLine($" Error renaming txt files to .bak: {ex.Message}");
}
}
private static void WritePostToFile(StreamWriter writer, TTPostRecord post)
{
var startColumns = new[] { "Post ID", "Date", "Post URL", "Slug", "Reblog Key", "Reblog URL", "Reblog Name", "Title", "Body" };
var endColumns = new[] { "Tags", "Downloaded Files" };
var columns = new Dictionary<string, string>();
if (!string.IsNullOrWhiteSpace(post.PostId)) columns["Post ID"] = post.PostId;
if (!string.IsNullOrWhiteSpace(post.Date)) columns["Date"] = post.Date!;
if (!string.IsNullOrWhiteSpace(post.PostUrl)) columns["Post URL"] = post.PostUrl!;
if (!string.IsNullOrWhiteSpace(post.Slug)) columns["Slug"] = post.Slug!;
if (!string.IsNullOrWhiteSpace(post.ReblogKey)) columns["Reblog Key"] = post.ReblogKey!;
if (!string.IsNullOrWhiteSpace(post.ReblogUrl)) columns["Reblog URL"] = post.ReblogUrl!;
if (!string.IsNullOrWhiteSpace(post.ReblogName)) columns["Reblog Name"] = post.ReblogName!;
if (!string.IsNullOrWhiteSpace(post.Title)) columns["Title"] = post.Title!;
if (!string.IsNullOrWhiteSpace(post.Body)) columns["Body"] = post.Body!;
if (!string.IsNullOrWhiteSpace(post.HasImage)) columns["Has Image"] = post.HasImage!;
if (!string.IsNullOrWhiteSpace(post.Summary)) columns["Summary"] = post.Summary!;
if (!string.IsNullOrWhiteSpace(post.Quote)) columns["Quote"] = post.Quote!;
if (!string.IsNullOrWhiteSpace(post.Link)) columns["Link"] = post.Link!;
if (!string.IsNullOrWhiteSpace(post.PhotoUrl)) columns["Photo URL"] = post.PhotoUrl!;
if (!string.IsNullOrWhiteSpace(post.PhotoCaption)) columns["Photo Caption"] = post.PhotoCaption!;
if (!string.IsNullOrWhiteSpace(post.AudioCaption)) columns["Audio Caption"] = post.AudioCaption!;
if (!string.IsNullOrWhiteSpace(post.Question)) columns["Question"] = post.Question!;
if (!string.IsNullOrWhiteSpace(post.Answer)) columns["Answer"] = post.Answer!;
if (!string.IsNullOrWhiteSpace(post.Tags)) columns["Tags"] = post.Tags!;
if (!string.IsNullOrWhiteSpace(post.DownloadedFiles)) columns["Downloaded Files"] = post.DownloadedFiles!;
foreach (var col in startColumns)
{
if (columns.ContainsKey(col))
{
writer.WriteLine($"{col}: {columns[col]}");
columns.Remove(col);
}
}
var remaining = columns.Keys.Where(k => !endColumns.Contains(k)).OrderBy(k => k).ToList();
foreach (var col in remaining)
writer.WriteLine($"{col}: {columns[col]}");
foreach (var col in endColumns)
{
if (columns.ContainsKey(col))
writer.WriteLine($"{col}: {columns[col]}");
}
}
}
}
+334 -84
View File
@@ -5,7 +5,6 @@ using Microsoft.Extensions.Configuration;
using System.Configuration; using System.Configuration;
using System.Threading; using System.Threading;
using Microsoft.Extensions.Diagnostics.Latency; using Microsoft.Extensions.Diagnostics.Latency;
using static System.Runtime.InteropServices.JavaScript.JSType;
using System.Text.RegularExpressions; using System.Text.RegularExpressions;
namespace URLNotesGrabberCORE namespace URLNotesGrabberCORE
@@ -13,12 +12,27 @@ namespace URLNotesGrabberCORE
internal class Program internal class Program
{ {
static void Main(string[] args) static int Main(string[] args)
{ {
// Reset console color on exit (including Ctrl+C) // Reset console color on exit (including Ctrl+C)
Console.CancelKeyPress += (s, e) => Console.ResetColor(); Console.CancelKeyPress += (s, e) => Console.ResetColor();
AppDomain.CurrentDomain.ProcessExit += (s, e) => Console.ResetColor(); AppDomain.CurrentDomain.ProcessExit += (s, e) => Console.ResetColor();
try
{
return Run(args);
}
catch (Exception ex)
{
// Turn configuration errors / unguarded indexers into a quiet, deterministic exit code.
Console.Error.WriteLine(ex.Message);
return 1;
}
}
static int Run(string[] args)
{
int exitCode = 0;
IConfiguration config = new ConfigurationBuilder() IConfiguration config = new ConfigurationBuilder()
.SetBasePath(Directory.GetCurrentDirectory()) .SetBasePath(Directory.GetCurrentDirectory())
.AddJsonFile("appsettings.json", optional: true, reloadOnChange: false) .AddJsonFile("appsettings.json", optional: true, reloadOnChange: false)
@@ -29,27 +43,38 @@ namespace URLNotesGrabberCORE
string apiSectionName = "TumblrApi"; string apiSectionName = "TumblrApi";
bool apiExplicitlySet = false; bool apiExplicitlySet = false;
string startFromBlogName = string.Empty; string startFromBlogName = string.Empty;
bool forceIgnoreCooldown = false;
List<string> filteredArgs = new List<string>(); List<string> filteredArgs = new List<string>();
for (int i = 0; i < args.Length; i++) for (int i = 0; i < args.Length; i++)
{ {
if (string.Equals(args[i], "-api3", StringComparison.OrdinalIgnoreCase) || if (args[i] == "--")
string.Equals(args[i], "--api3", StringComparison.OrdinalIgnoreCase)) {
// POSIX end-of-options: everything after is a literal operand.
for (int j = i + 1; j < args.Length; j++) filteredArgs.Add(args[j]);
break;
}
if (string.Equals(args[i], "--force", StringComparison.OrdinalIgnoreCase))
{
forceIgnoreCooldown = true;
continue;
}
if (string.Equals(args[i], "--api3", StringComparison.OrdinalIgnoreCase))
{ {
apiSectionName = "TumblrApi3"; apiSectionName = "TumblrApi3";
apiExplicitlySet = true; apiExplicitlySet = true;
continue; continue;
} }
if (string.Equals(args[i], "-api4", StringComparison.OrdinalIgnoreCase) || if (string.Equals(args[i], "--api4", StringComparison.OrdinalIgnoreCase))
string.Equals(args[i], "--api4", StringComparison.OrdinalIgnoreCase))
{ {
apiSectionName = "TumblrApi4"; apiSectionName = "TumblrApi4";
apiExplicitlySet = true; apiExplicitlySet = true;
continue; continue;
} }
if (string.Equals(args[i], "-api", StringComparison.OrdinalIgnoreCase) || if (string.Equals(args[i], "--api", StringComparison.OrdinalIgnoreCase))
string.Equals(args[i], "--api", StringComparison.OrdinalIgnoreCase))
{ {
if (i + 1 < args.Length && !string.IsNullOrWhiteSpace(args[i + 1])) if (i + 1 < args.Length && !string.IsNullOrWhiteSpace(args[i + 1]))
{ {
@@ -59,13 +84,12 @@ namespace URLNotesGrabberCORE
} }
else else
{ {
Console.WriteLine("--Missing API section after -api/--api. Using default TumblrApi.--"); Console.WriteLine("--Missing API section after --api. Using default TumblrApi.--");
} }
continue; continue;
} }
if (string.Equals(args[i], "-start", StringComparison.OrdinalIgnoreCase) || if (string.Equals(args[i], "--start", StringComparison.OrdinalIgnoreCase))
string.Equals(args[i], "--start", StringComparison.OrdinalIgnoreCase))
{ {
if (i + 1 < args.Length && !string.IsNullOrWhiteSpace(args[i + 1])) if (i + 1 < args.Length && !string.IsNullOrWhiteSpace(args[i + 1]))
{ {
@@ -74,7 +98,7 @@ namespace URLNotesGrabberCORE
} }
else else
{ {
Console.WriteLine("--Missing blog name after -start/--start. Ignoring.--"); Console.WriteLine("--Missing blog name after --start. Ignoring.--");
} }
continue; continue;
} }
@@ -115,7 +139,10 @@ namespace URLNotesGrabberCORE
Console.SetOut(dualLogger); Console.SetOut(dualLogger);
} }
List<string> contains = settings.GetValue<string>("ContainsList").Split(',').ToList(); string? containsListSetting = settings.GetValue<string>("ContainsList");
if (string.IsNullOrEmpty(containsListSetting))
throw new InvalidOperationException("ContainsList is not configured in appsettings.json");
List<string> contains = containsListSetting.Split(',').ToList();
bool logTraversalRecordImports = settings.GetValue("LogTraversalRecordImports", false); bool logTraversalRecordImports = settings.GetValue("LogTraversalRecordImports", false);
if (args.Length == 0) //Traverse folder structure to add posts and thus blogs to DB if (args.Length == 0) //Traverse folder structure to add posts and thus blogs to DB
@@ -124,10 +151,12 @@ namespace URLNotesGrabberCORE
try try
{ {
DataAccess.EnableImportModePragmas(); DataAccess.EnableImportModePragmas();
DataAccess.BeginImportSession();
TraverseDirectory(settings.GetValue<string>("PathInput"), settings.GetValue<string>("PathOutputBlogs"), contains, ref postsAdded, startFromBlogName: startFromBlogName, logRecordImports: logTraversalRecordImports); TraverseDirectory(settings.GetValue<string>("PathInput"), settings.GetValue<string>("PathOutputBlogs"), contains, ref postsAdded, startFromBlogName: startFromBlogName, logRecordImports: logTraversalRecordImports);
} }
finally finally
{ {
DataAccess.EndImportSession();
DataAccess.RestoreImportModePragmas(); DataAccess.RestoreImportModePragmas();
} }
Console.WriteLine($"Total posts added: {postsAdded}"); Console.WriteLine($"Total posts added: {postsAdded}");
@@ -137,41 +166,23 @@ namespace URLNotesGrabberCORE
switch (args[0]) switch (args[0])
{ {
case "-?": case "-?":
Console.WriteLine("\t Parse .txt files to find blogs"); case "-h":
case "--help":
Console.WriteLine("-?\t Usage help"); PrintHelp();
Console.WriteLine("-parse\t Parse .txt files with specified blogname");
Console.WriteLine("-test\t Calls API for given blogname and postID");
Console.WriteLine("-posts\t For each Post in DB, write blogname to file");
Console.WriteLine("-blogs\t For each Blog in DB, write blogname to file");
Console.WriteLine("-collect\t For each Post in DB, hit API to collect Notes. Optional datetime parameter to filter by NotesGatheredDateTime");
Console.WriteLine("-blogsR\t For each Note that is a REPLY, write blogname to file ");
Console.WriteLine("-blogsO\t For each Blog in DB, write blogname to file, but limit via a passed start and stop range ");
Console.WriteLine("-replies\t Fetch and update missing reply text for all replies in database");
Console.WriteLine("-likes\t Fetch likes for all blogs needing it (LikesPulled=0), or a specific blog via param");
Console.WriteLine("-urldump\t Scan all posts' text columns and extract suspected URLs to configured file");
Console.WriteLine("-api3\t Use TumblrApi3 settings from appsettings.json");
Console.WriteLine("-api4\t Use TumblrApi4 settings from appsettings.json");
Console.WriteLine("-start [blogname]\t Start traversal alphabetically at this blog name");
Console.WriteLine("-api [section]\t Use a specific API settings section from appsettings.json (e.g. TumblrApi3)");
break; break;
case "-parse": case "-V":
case "--version":
Console.WriteLine(System.Reflection.Assembly.GetExecutingAssembly().GetName().Version?.ToString() ?? "unknown");
break;
case "--parse":
if (args.Length < 2)
{
Console.WriteLine("Usage: --parse <blogname>");
exitCode = 2;
break;
}
string blogNameToParse = args[1]; string blogNameToParse = args[1];
int postsAdded = 0; int postsAdded = 0;
try try
@@ -186,29 +197,31 @@ namespace URLNotesGrabberCORE
Console.WriteLine($"Total posts added: {postsAdded}"); Console.WriteLine($"Total posts added: {postsAdded}");
break; break;
case "-test": case "--test":
Console.WriteLine("Test command not implemented"); Console.WriteLine("Test command not implemented");
break; break;
case "-post": case "--post":
TraverseDirectoryForCorruption(settings.GetValue<string>("PathInput"), settings.GetValue<string>("PathOutputBlogs"), contains); TraverseDirectoryForCorruption(settings.GetValue<string>("PathInput"), settings.GetValue<string>("PathOutputBlogs"), contains);
break; break;
case "-posts": //write post's blogs to file case "--posts": //write post's blogs to file
WritePostBlogsToFile(settings.GetValue<string>("PathOutputPosts")); WritePostBlogsToFile(settings.GetValue<string>("PathOutputPosts"));
break; break;
case "-blogs": //write blogs to file case "--blogs": //write blogs to file
WriteBlogsToFile(settings.GetValue<string>("PathOutputBlogs")); WriteBlogsToFile(settings.GetValue<string>("PathOutputBlogs"));
break; break;
case "-collect": //collect notes from all posts case "--collect": //collect notes from all posts
bool withoutNotesOnly = true; bool withoutNotesOnly = true;
DateTime? beforeDate = DateTime.Now; DateTime? beforeDate = DateTime.Now;
bool explicitDateSupplied = false;
if (args.Length < 2) if (args.Length < 2)
{ {
Console.WriteLine("--Expected WITHOUTNOTESONLY (0, 1) [OPTIONAL: BEFOREDATE]--"); Console.WriteLine("--Expected WITHOUTNOTESONLY (0, 1) [OPTIONAL: BEFOREDATE]--");
exitCode = 2;
break; break;
} }
@@ -235,26 +248,50 @@ namespace URLNotesGrabberCORE
if (DateTime.TryParse(args[2], out DateTime parsedDate)) if (DateTime.TryParse(args[2], out DateTime parsedDate))
{ {
beforeDate = parsedDate; beforeDate = parsedDate;
explicitDateSupplied = true;
Console.WriteLine($"Filter: Collecting notes for posts with NotesGatheredDateTime < {beforeDate}"); Console.WriteLine($"Filter: Collecting notes for posts with NotesGatheredDateTime < {beforeDate}");
} }
else else
{ {
Console.WriteLine($"ERROR: Invalid date format '{args[2]}'"); Console.WriteLine($"ERROR: Invalid date format '{args[2]}'");
exitCode = 2;
break; break;
} }
} }
CollectNotes(settings.GetValue<string>("PathOutput"), withoutNotesOnly, beforeDate).GetAwaiter().GetResult(); // Mode 0 (full re-check) with no explicit date is a *managed* run: freeze the cutoff and
// persist it so an interrupted run resumes against the same cutoff and a completed run stops
// instead of restarting. Mode 1 and explicit-date runs keep their existing behavior.
bool managedCollectRun = false;
if (!withoutNotesOnly && !explicitDateSupplied)
{
DataAccess.EnsureCollectRunStateTableExists();
var runState = DataAccess.GetCollectRunState();
if (runState != null && !runState.Value.complete)
{
beforeDate = DateTimeOffset.FromUnixTimeSeconds(runState.Value.cutoff).LocalDateTime;
Console.WriteLine($"Resuming interrupted full re-check (cutoff = {beforeDate})");
}
else
{
beforeDate = DateTime.Now;
DataAccess.BeginCollectRun(new DateTimeOffset(beforeDate.Value).ToUnixTimeSeconds());
Console.WriteLine($"Starting new full re-check run (cutoff = {beforeDate})");
}
managedCollectRun = true;
}
CollectNotes(settings.GetValue<string>("PathOutput"), withoutNotesOnly, beforeDate, managedCollectRun).GetAwaiter().GetResult();
break; break;
case "-blogsR": //collect notes from all posts case "--blogsR": //collect notes from all posts
WriteBlogsToFile(settings.GetValue<string>("PathOutputBlogs"), true); WriteBlogsToFile(settings.GetValue<string>("PathOutputBlogs"), true);
break; break;
case "-blogsO": //collect notes from all posts case "--blogsO": //collect notes from all posts
int from = 1, to = 999999, top = 100; int from = 1, to = 999999, top = 100;
if (args[1] is not null && args[2] is not null && args[3] is not null) if (args.Length >= 4 && args[1] is not null && args[2] is not null && args[3] is not null)
{ {
from = int.Parse(args[1]); from = int.Parse(args[1]);
to = int.Parse(args[2]); to = int.Parse(args[2]);
@@ -263,14 +300,16 @@ namespace URLNotesGrabberCORE
else else
{ {
Console.WriteLine("--Expected FROM TO--"); Console.WriteLine("--Expected FROM TO--");
exitCode = 2;
break;
} }
WriteBlogsToFile(settings.GetValue<string>("PathOutputBlogs"), false, from, to, top); WriteBlogsToFile(settings.GetValue<string>("PathOutputBlogs"), false, from, to, top);
break; break;
case "-bop": //collect notes from all posts case "--bop": //collect notes from all posts
from = 1; to = 999999; top = 100; from = 1; to = 999999; top = 100;
if (args[1] is not null && args[2] is not null && args[3] is not null) if (args.Length >= 4 && args[1] is not null && args[2] is not null && args[3] is not null)
{ {
from = int.Parse(args[1]); from = int.Parse(args[1]);
to = int.Parse(args[2]); to = int.Parse(args[2]);
@@ -279,25 +318,70 @@ namespace URLNotesGrabberCORE
else else
{ {
Console.WriteLine("--Expected FROM TO--"); Console.WriteLine("--Expected FROM TO--");
exitCode = 2;
break;
} }
WriteBlogsToFileAll(settings.GetValue<string>("PathOutputBlogs"), false, from, to, top); WriteBlogsToFileAll(settings.GetValue<string>("PathOutputBlogs"), false, from, to, top);
break; break;
case "-replies": //update reply text case "--replies": //update reply text
CollectMissingReplyText().GetAwaiter().GetResult(); CollectMissingReplyText().GetAwaiter().GetResult();
break; break;
case "-likes": case "--likes":
string likeBlog = args.Length > 1 ? args[1] : null; string likeBlog = args.Length > 1 ? args[1] : null;
CollectLikes(likeBlog, contains).GetAwaiter().GetResult(); int cooldownDays = settings.GetValue("LikesRefreshCooldownDays", 7);
CollectLikes(likeBlog, contains, cooldownDays, forceIgnoreCooldown).GetAwaiter().GetResult();
break; break;
case "-urldump": case "--urldump":
DumpUrls(settings.GetValue<string>("PathOutputUrls")); DumpUrls(settings.GetValue<string>("PathOutputUrls"));
break; break;
case "--ingest":
exitCode = IngestMode.Run(config, args.Skip(1).ToArray());
break;
case "--output":
exitCode = OutputMode.Run(config);
break;
case "--revert":
exitCode = RevertMode.Run(config, args.Length > 1 ? args[1] : null);
break;
case "--correct":
{
bool applyChanges = args.Skip(1).Any(a => string.Equals(a, "--apply", StringComparison.OrdinalIgnoreCase));
var correctArgs = args.Skip(1)
.Where(a => !string.Equals(a, "--apply", StringComparison.OrdinalIgnoreCase))
.ToArray();
exitCode = CorrectMode.Run(config, correctArgs, applyChanges);
break;
}
case "--updatepaths":
{
string rootPath = args.Length > 1 ? args[1] : (settings.GetValue<string>("PathTTRoot") ?? settings.GetValue<string>("PathInput") ?? string.Empty);
exitCode = UpdateBlogPathsRunner.Run(rootPath);
break;
}
case "--importposts":
{
if (args.Length < 2)
{
Console.WriteLine("Usage: --importposts <path-to-legacy-posts.db>");
exitCode = 2;
break;
}
exitCode = LegacyPostsDbImporter.Run(args[1]);
break;
}
default: default:
Console.WriteLine("** Unknown Command ** " + args[0]); Console.WriteLine("** Unknown Command ** " + args[0]);
exitCode = 2;
break; break;
} }
} }
@@ -306,6 +390,69 @@ namespace URLNotesGrabberCORE
System.Console.WriteLine("<fin>:/"); System.Console.WriteLine("<fin>:/");
//System.Console.ReadKey(); //System.Console.ReadKey();
return exitCode;
}
static void PrintHelp()
{
Console.WriteLine("\t Parse .txt files to find blogs");
Console.WriteLine("-?, -h, --help\t Usage help");
Console.WriteLine("-V, --version\t Print the application version");
Console.WriteLine("--\t End of options: treat every following token as a literal operand");
Console.WriteLine("--parse\t Parse .txt files with specified blogname");
Console.WriteLine("--test\t Calls API for given blogname and postID");
Console.WriteLine("--post\t Traverse the input directory tree checking .txt files for corruption");
Console.WriteLine("--posts\t For each Post in DB, write blogname to file");
Console.WriteLine("--blogs\t For each Blog in DB, write blogname to file");
Console.WriteLine("--collect [0|1] [datetime]\t Collect Notes from API. 1=only posts without notes. 0=full re-check of all posts: a single resumable pass (interrupt & relaunch to resume; stops when complete, retrigger for a new pass). Optional datetime overrides the cutoff and runs as a one-off (bypasses resume tracking).");
Console.WriteLine("--blogsR\t For each Note that is a REPLY, write blogname to file ");
Console.WriteLine("--blogsO\t For each Blog in DB, write blogname to file, but limit via a passed start and stop range ");
Console.WriteLine("--bop [from] [to] [top]\t Write ALL blog names to file, limited by FROM TO TOP range arguments");
Console.WriteLine("--replies\t Fetch and update missing reply text for all replies in database");
Console.WriteLine("--likes\t Fetch likes: initial backfill for new blogs, incremental refresh for blogs past cooldown. Optional blog name forces single-blog run.");
Console.WriteLine("--force\t (with --likes) Ignore cooldown and refresh every fully-backfilled blog");
Console.WriteLine("--urldump\t Scan all posts' text columns and extract suspected URLs to configured file");
Console.WriteLine("--api3\t Use TumblrApi3 settings from appsettings.json");
Console.WriteLine("--api4\t Use TumblrApi4 settings from appsettings.json");
Console.WriteLine("--start [blogname]\t Start traversal alphabetically at this blog name");
Console.WriteLine("--api [section]\t Use a specific API settings section from appsettings.json (e.g. TumblrApi3)");
Console.WriteLine("--ingest [blogname]\t Ingest Tumblr .txt exports from appSettings:PathTTRoot into TL.db (all blogs, or single blog if name given)");
Console.WriteLine("--output\t Export posts from TL.db back to .txt files in each blog's TTFolderPath");
Console.WriteLine("--revert [blogname]\t Recursively scan the PathInput tree and restore *.bak back to *.txt (current .txt saved as next-free .bkN); optional blogname filters by path substring");
Console.WriteLine("--correct [bakPath]\t Dry-run: report multi-line field updates available from a BAK directory");
Console.WriteLine("--correct --apply [bakPath]\t Apply BAK-file corrections to matching posts (prompts yes/no)");
Console.WriteLine("--updatepaths [rootPath]\t Read .tumblr/.tmblrpriv metadata from <root>\\Index and set Blogs.TTFolderPath");
Console.WriteLine("--importposts [path-to-posts.db]\t One-time migration: copy legacy ThreeTxtFileHelper posts.db rows into TL.db");
Console.WriteLine();
Console.WriteLine("Exit status: 0 = success; 1 = unexpected error; 2 = usage error (unknown command or bad/missing arguments)");
} }
static void WritePostBlogsToFile(string outPath) static void WritePostBlogsToFile(string outPath)
@@ -679,14 +826,14 @@ namespace URLNotesGrabberCORE
} }
} }
static async Task CollectLikes(string specificBlog, List<string> contains) static async Task CollectLikes(string specificBlog, List<string> contains, int cooldownDays = 7, bool ignoreCooldown = false)
{ {
try try
{ {
DataAccess.EnsureBlogsLikesColumnsExist(); DataAccess.EnsureBlogsLikesColumnsExist();
Console.WriteLine("Starting collection of likes..."); Console.WriteLine($"Starting collection of likes... (cooldown {cooldownDays}d, ignoreCooldown={ignoreCooldown})");
var blogsToProcess = DataAccess.GetBlogsForLikes(specificBlog); var blogsToProcess = DataAccess.GetBlogsForLikes(specificBlog, cooldownDays, ignoreCooldown);
if (blogsToProcess.Count == 0) if (blogsToProcess.Count == 0)
{ {
@@ -694,7 +841,9 @@ namespace URLNotesGrabberCORE
return; return;
} }
Console.WriteLine($"Found {blogsToProcess.Count} blogs to process likes."); int backfillCount = blogsToProcess.Count(b => b.Item2 == 0);
int refreshCount = blogsToProcess.Count(b => b.Item2 == 1);
Console.WriteLine($"Found {blogsToProcess.Count} blogs to process likes ({backfillCount} backfill, {refreshCount} refresh).");
RateLimiter limiter = new SlidingWindowRateLimiter(new SlidingWindowRateLimiterOptions RateLimiter limiter = new SlidingWindowRateLimiter(new SlidingWindowRateLimiterOptions
{ {
@@ -706,18 +855,31 @@ namespace URLNotesGrabberCORE
AutoReplenishment = true AutoReplenishment = true
}); });
foreach (var blogInfo in blogsToProcess) for (int i = 0; i < blogsToProcess.Count; i++)
{ {
var blogInfo = blogsToProcess[i];
int remaining = blogsToProcess.Count - i - 1;
string blogName = blogInfo.Item1; string blogName = blogInfo.Item1;
int likesPulled = blogInfo.Item2; int likesPulled = blogInfo.Item2;
long cursor = blogInfo.Item3; long cursor = blogInfo.Item3;
long storedNewestTs = blogInfo.Item4;
bool isRefresh = likesPulled == 1;
long parsedForBlog = 0; long parsedForBlog = 0;
long matchedForBlog = 0; long matchedForBlog = 0;
int likedCountForBlog = 0; int likedCountForBlog = 0;
long observedMaxLikedTs = storedNewestTs;
int newInsertedInRefresh = 0;
Console.WriteLine($"Processing likes for blog: {blogName} | Cursor: {cursor}"); // Refresh always starts from the top (newest) and walks backward until it crosses
// the stored high-water mark. Backfill resumes from its last persisted cursor.
if (isRefresh) cursor = 0;
string mode = isRefresh ? "REFRESH" : "BACKFILL";
Console.WriteLine($"[{remaining} remaining] Processing likes for blog: {blogName} | Mode: {mode} | Cursor: {cursor} | HighWaterMark: {storedNewestTs}");
bool hasMoreLikes = true; bool hasMoreLikes = true;
bool isFirstPage = true;
while (hasMoreLikes) while (hasMoreLikes)
{ {
@@ -746,14 +908,23 @@ namespace URLNotesGrabberCORE
if (response?.statusCode == "NotFound" || (response?.meta != null && response.meta.status == 404)) if (response?.statusCode == "NotFound" || (response?.meta != null && response.meta.status == 404))
{ {
Console.WriteLine($"API returned 404 Not Found for {blogName} Likes"); Console.WriteLine($"API returned 404 Not Found for {blogName} Likes");
DataAccess.UpdateBlogLikesStatus(blogName, 1, cursor); if (isRefresh)
DataAccess.UpdateBlogLikesRefreshStatus(blogName, observedMaxLikedTs, newInsertedInRefresh);
else
DataAccess.UpdateBlogLikesStatus(blogName, 1, cursor);
break; break;
} }
if (response?.response?.liked_posts == null || response.response.liked_posts.Count == 0) if (response?.response?.liked_posts == null || response.response.liked_posts.Count == 0)
{ {
Console.WriteLine($"[Likes] No more likes found for {blogName}. Marking complete."); Console.WriteLine($"[Likes] No more likes found for {blogName}. Marking complete.");
DataAccess.UpdateBlogLikesStatus(blogName, 1, cursor); // Done parsing if (isRefresh)
DataAccess.UpdateBlogLikesRefreshStatus(blogName, observedMaxLikedTs, newInsertedInRefresh);
else
{
DataAccess.UpdateBlogLikesStatus(blogName, 1, cursor);
DataAccess.UpdateBlogLikesNewestTimestamp(blogName, observedMaxLikedTs);
}
hasMoreLikes = false; hasMoreLikes = false;
break; break;
} }
@@ -761,7 +932,9 @@ namespace URLNotesGrabberCORE
if (response.response.liked_count > 0) if (response.response.liked_count > 0)
likedCountForBlog = response.response.liked_count; likedCountForBlog = response.response.liked_count;
Console.WriteLine($"[Likes] Fetched {response.response.liked_posts.Count} likes for {blogName}"); Console.WriteLine($"[Likes] [Fetched {response.response.liked_posts.Count}] {blogName}");
bool crossedHighWaterMark = false;
foreach (var post in response.response.liked_posts) foreach (var post in response.response.liked_posts)
{ {
@@ -770,6 +943,20 @@ namespace URLNotesGrabberCORE
long postID = 0; long postID = 0;
try { postID = Convert.ToInt64(post.id); } catch { continue; } try { postID = Convert.ToInt64(post.id); } catch { continue; }
// liked_timestamp is when the user liked the post (matches the `before` cursor semantics).
// It's the only reliable field for the refresh stop condition.
long likedTs = 0;
try { likedTs = Convert.ToInt64(post.liked_timestamp); } catch { }
if (isRefresh && likedTs > 0 && storedNewestTs > 0 && likedTs <= storedNewestTs)
{
Console.WriteLine($"[Likes] Reached high-water mark for {blogName} at liked_timestamp={likedTs} (<= stored {storedNewestTs}). Stopping refresh.");
crossedHighWaterMark = true;
break;
}
if (likedTs > observedMaxLikedTs) observedMaxLikedTs = likedTs;
string authorBlog = post.blog_name?.ToString() ?? "."; string authorBlog = post.blog_name?.ToString() ?? ".";
string postURL = post.post_url?.ToString() ?? "."; string postURL = post.post_url?.ToString() ?? ".";
string date = post.date?.ToString() ?? "."; string date = post.date?.ToString() ?? ".";
@@ -858,20 +1045,37 @@ namespace URLNotesGrabberCORE
} }
} }
if (shouldInsert) if (shouldInsert)
{ {
matchedForBlog++; matchedForBlog++;
Console.WriteLine($"[Likes] Match | Author: {authorBlog} | PostID: {postID} | Field: {matchedFieldName}"); if (isRefresh) newInsertedInRefresh++;
Console.WriteLine($"[Likes] [Matched] {blogName} | Author: {authorBlog} | PostID: {postID} | Field: {matchedFieldName}");
await Task.Delay(3000); await Task.Delay(3000);
DataAccess.AddPost(authorBlog, postID, reblogURL, date, postURL, slug, reblogKey, DataAccess.AddPost(authorBlog, postID, reblogURL, date, postURL, slug, reblogKey,
reblogName, summary, quote, body, tags, link, photoURL, reblogName, summary, quote, body, tags, link, photoURL,
photoCaption, downloadedFiles, audioCaption, question, answer, photoCaption, downloadedFiles, audioCaption, question, answer,
title, hasImage, true, rootBlogName: rootBlogName, rootURL: rootURL); title, hasImage, true, rootBlogName: rootBlogName, rootURL: rootURL);
} }
} }
WriteLikesTotalsLine(blogName, "Running Totals", parsedForBlog, matchedForBlog, likedCountForBlog); WriteLikesTotalsLine(blogName, "Running Totals", parsedForBlog, matchedForBlog, likedCountForBlog);
// Branch B: refresh terminates as soon as we crossed the high-water mark.
if (isRefresh && crossedHighWaterMark)
{
DataAccess.UpdateBlogLikesRefreshStatus(blogName, observedMaxLikedTs, newInsertedInRefresh);
hasMoreLikes = false;
break;
}
// Branch A: on the very first page, capture & persist the newest liked_timestamp
// so subsequent refresh runs (after backfill completes) have a stopping point.
if (!isRefresh && isFirstPage && observedMaxLikedTs > 0)
{
DataAccess.UpdateBlogLikesNewestTimestamp(blogName, observedMaxLikedTs);
}
isFirstPage = false;
// Determine the next BeforeCursor. // Determine the next BeforeCursor.
long nextCursor = 0; long nextCursor = 0;
if (response.response._links?.next?.query_params != null) if (response.response._links?.next?.query_params != null)
@@ -879,9 +1083,12 @@ namespace URLNotesGrabberCORE
long.TryParse(response.response._links.next.query_params.before ?? "0", out nextCursor); long.TryParse(response.response._links.next.query_params.before ?? "0", out nextCursor);
} }
// Persist cursor progress after every page so resume is always up-to-date if (!isRefresh)
long cursorToPersist = nextCursor > 0 ? nextCursor : cursor; {
DataAccess.UpdateBlogLikesStatus(blogName, 0, cursorToPersist); // Persist cursor progress after every page so resume is always up-to-date
long cursorToPersist = nextCursor > 0 ? nextCursor : cursor;
DataAccess.UpdateBlogLikesStatus(blogName, 0, cursorToPersist);
}
if (nextCursor > 0) if (nextCursor > 0)
{ {
@@ -891,17 +1098,29 @@ namespace URLNotesGrabberCORE
else else
{ {
Console.WriteLine($"[Likes] No further pagination items. Done with {blogName}."); Console.WriteLine($"[Likes] No further pagination items. Done with {blogName}.");
DataAccess.UpdateBlogLikesStatus(blogName, 1, cursorToPersist); // Mark as completely pulled if (isRefresh)
{
DataAccess.UpdateBlogLikesRefreshStatus(blogName, observedMaxLikedTs, newInsertedInRefresh);
}
else
{
long cursorToPersist = nextCursor > 0 ? nextCursor : cursor;
DataAccess.UpdateBlogLikesStatus(blogName, 1, cursorToPersist);
DataAccess.UpdateBlogLikesNewestTimestamp(blogName, observedMaxLikedTs);
}
hasMoreLikes = false; hasMoreLikes = false;
} }
Console.WriteLine($"[{remaining} remaining] Done with {blogName}");
await Task.Delay(1000); // 1-second delay between pages await Task.Delay(1000); // 1-second delay between pages
} }
if (isRefresh)
Console.WriteLine($"[Likes] {blogName} refresh complete | New inserted: {newInsertedInRefresh} | New HighWaterMark: {observedMaxLikedTs}");
WriteLikesTotalsLine(blogName, "Final Totals", parsedForBlog, matchedForBlog, likedCountForBlog); WriteLikesTotalsLine(blogName, "Final Totals", parsedForBlog, matchedForBlog, likedCountForBlog);
} }
Console.WriteLine("Likes collection complete."); Console.WriteLine("[Likes] [Likes collection complete]");
} }
catch (Exception ex) catch (Exception ex)
{ {
@@ -1049,10 +1268,15 @@ namespace URLNotesGrabberCORE
return "UNKNOWN"; return "UNKNOWN";
} }
static async Task CollectNotes(string outPath, bool withoutNotesOnly = true, DateTime? beforeDate = null) static async Task CollectNotes(string outPath, bool withoutNotesOnly = true, DateTime? beforeDate = null, bool managedRun = false)
{ {
List<Tuple<string, long, long, long>> posts = DataAccess.GetPosts(withoutNotesOnly, beforeDate); List<Tuple<string, long, long, long>> posts = DataAccess.GetPosts(withoutNotesOnly, beforeDate);
// Posts attempted (with a definitive, non-throttle result) during *this* process. Guarantees a single
// attempt pass: once every remaining post has been attempted, the loop stops instead of spinning on a
// post that keeps returning FAILURE/UNKNOWN. Successful/NotFound posts drop out via the DB filter anyway.
HashSet<(string, long)> attempted = new HashSet<(string, long)>();
RateLimiter limiter = new SlidingWindowRateLimiter(new SlidingWindowRateLimiterOptions RateLimiter limiter = new SlidingWindowRateLimiter(new SlidingWindowRateLimiterOptions
{ {
PermitLimit = 300, PermitLimit = 300,
@@ -1069,9 +1293,17 @@ namespace URLNotesGrabberCORE
{ {
while (posts.Count > 0) while (posts.Count > 0)
{ {
// First post not yet attempted this process. If all remaining have been attempted, the pass
// is done (the stragglers returned FAILURE/UNKNOWN) — stop rather than loop forever.
var post = posts.FirstOrDefault(p => !attempted.Contains((p.Item1, p.Item2)));
if (post == null)
{
Console.WriteLine("All remaining posts have been attempted this run; ending pass.");
break;
}
ApiKeyPool.SleepUntilAnyAvailable(30); ApiKeyPool.SleepUntilAnyAvailable(30);
var post = posts[0]; // Process the first post in the list
string status; string status;
using RateLimitLease lease = limiter.AttemptAcquire(1); using RateLimitLease lease = limiter.AttemptAcquire(1);
@@ -1083,20 +1315,31 @@ namespace URLNotesGrabberCORE
else else
{ {
Console.WriteLine("!@@@@@@@ - Rate Limited Exceeded: No Lease Available"); Console.WriteLine("!@@@@@@@ - Rate Limited Exceeded: No Lease Available");
return; return; // throttle: abort without completing the run so a later launch resumes
} }
if (status == "Success") if (status == "Success")
{ {
attempted.Add((post.Item1, post.Item2));
DataAccess.UpdatePostMarkNotesCollected(post.Item1, post.Item2); DataAccess.UpdatePostMarkNotesCollected(post.Item1, post.Item2);
} }
else if (status == "NotFound") else if (status == "NotFound")
{ {
attempted.Add((post.Item1, post.Item2));
Console.WriteLine("GrabNotes Result: NotFound"); Console.WriteLine("GrabNotes Result: NotFound");
DataAccess.UpdatePostMarkNotFound(post.Item1, post.Item2); DataAccess.UpdatePostMarkNotFound(post.Item1, post.Item2);
} }
else if (status == "TooManyRequests")
{
// Throttle, not a real per-post failure: don't consume this post's single attempt.
// Abort the pass without completing so a later launch resumes against the same cutoff.
Console.WriteLine("GrabNotes Result: TooManyRequests - pausing run; relaunch to resume.");
return;
}
else else
{ {
// FAILURE / UNKNOWN: count as attempted so the pass can finish instead of retrying forever.
attempted.Add((post.Item1, post.Item2));
Console.WriteLine("GrabNotes Result: " + status); Console.WriteLine("GrabNotes Result: " + status);
} }
@@ -1104,6 +1347,13 @@ namespace URLNotesGrabberCORE
posts = DataAccess.GetPosts(withoutNotesOnly, beforeDate); posts = DataAccess.GetPosts(withoutNotesOnly, beforeDate);
} }
} }
// Reached only when the pass finished naturally (worklist drained or all stragglers attempted).
if (managedRun)
{
DataAccess.CompleteCollectRun();
Console.WriteLine("Full re-check run complete.");
}
} }
catch (Exception ex) catch (Exception ex)
{ {
@@ -2,7 +2,7 @@
"profiles": { "profiles": {
"URLNotesGrabberCORE": { "URLNotesGrabberCORE": {
"commandName": "Project", "commandName": "Project",
"commandLineArgs": "-collect 1 -api4" "commandLineArgs": "--collect 1 --api4"
} }
} }
} }
+110
View File
@@ -0,0 +1,110 @@
using Microsoft.Extensions.Configuration;
namespace URLNotesGrabberCORE
{
// Inverse of OutputMode. Recursively walks the PathInput tree (the same directory tree the
// no-parameter run uses) and restores every *.bak back to its *.txt, first preserving the
// current *.txt as the next-free *.bkN. Consumes the *.bak (File.Move). Filesystem-only;
// does not read the DB. An optional blogname argument filters by path substring.
public static class RevertMode
{
public static int Run(IConfiguration config, string? blogFilter = null)
{
string? root = config["appSettings:PathInput"];
if (string.IsNullOrWhiteSpace(root) || !Directory.Exists(root))
{
Console.WriteLine($"PathInput is not set or does not exist: '{root}'. Nothing to revert.");
return 0;
}
Console.WriteLine($"Searching for .bak files under: {root}");
// Recursively collect every *.bak, optionally filtered by path substring (blogname).
var bakFiles = EnumerateBakFiles(root)
.Where(f => string.IsNullOrWhiteSpace(blogFilter)
|| f.IndexOf(blogFilter, StringComparison.OrdinalIgnoreCase) >= 0)
.ToList();
if (bakFiles.Count == 0)
{
Console.WriteLine("No .bak files found. Nothing to revert.");
return 0;
}
Console.Write($"WARNING: This will restore {bakFiles.Count} .bak file(s) over their .txt files. " +
$"Current .txt files are preserved as the next-free .bkN. Continue? (yes/no): ");
string? response = Console.ReadLine();
if (string.IsNullOrWhiteSpace(response) || !response.Equals("yes", StringComparison.OrdinalIgnoreCase))
{
Console.WriteLine("Operation cancelled.");
return 0;
}
int restored = 0, backedUp = 0;
foreach (var bakFile in bakFiles)
{
try
{
string txtPath = Path.ChangeExtension(bakFile, ".txt");
if (File.Exists(txtPath))
{
string bkPath = NextFreeBkPath(txtPath);
File.Move(txtPath, bkPath);
backedUp++;
Console.WriteLine($" Backed up {Path.GetFileName(txtPath)} -> {Path.GetFileName(bkPath)}");
}
File.Move(bakFile, txtPath);
restored++;
Console.WriteLine($" Restored {bakFile} -> {Path.GetFileName(txtPath)}");
}
catch (Exception ex)
{
Console.WriteLine($" Error reverting {bakFile}: {ex.Message}");
}
}
Console.WriteLine($"\nRevert mode complete. Restored {restored} file(s); backed up {backedUp} current .txt file(s).");
return 0;
}
// Recursively yields every *.bak path under root. Per-directory try/catch so an
// inaccessible folder doesn't abort the whole walk (mirrors TraverseDirectory).
private static IEnumerable<string> EnumerateBakFiles(string path)
{
string[] subDirs;
try { subDirs = Directory.GetDirectories(path); }
catch (Exception ex)
{
Console.WriteLine($" Skipping '{path}': {ex.Message}");
yield break;
}
foreach (var dir in subDirs)
foreach (var bak in EnumerateBakFiles(dir))
yield return bak;
string[] bakFiles;
try { bakFiles = Directory.GetFiles(path, "*.bak"); }
catch (Exception ex)
{
Console.WriteLine($" Skipping files in '{path}': {ex.Message}");
yield break;
}
foreach (var bak in bakFiles)
yield return bak;
}
// Returns the lowest unused .bkN path for a given .txt file (.bk1, .bk2, ...).
private static string NextFreeBkPath(string txtFile)
{
for (int n = 1; ; n++)
{
string candidate = Path.ChangeExtension(txtFile, $".bk{n}");
if (!File.Exists(candidate)) return candidate;
}
}
}
}
@@ -29,6 +29,9 @@
<CopyToOutputDirectory>Always</CopyToOutputDirectory> <CopyToOutputDirectory>Always</CopyToOutputDirectory>
</None> </None>
<None Update="TL.db"> <None Update="TL.db">
<CopyToOutputDirectory>Never</CopyToOutputDirectory>
</None>
<None Update="prefixes.txt">
<CopyToOutputDirectory>PreserveNewest</CopyToOutputDirectory> <CopyToOutputDirectory>PreserveNewest</CopyToOutputDirectory>
</None> </None>
</ItemGroup> </ItemGroup>
@@ -0,0 +1,70 @@
using System.Text.Json;
namespace URLNotesGrabberCORE
{
// Port of ThreeTxtFileHelper/UpdateBlogPaths.cs. Reads .tumblr / .tmblrpriv metadata
// files from a root\Index folder and populates Blogs.TTFolderPath in TL.db.
public static class UpdateBlogPathsRunner
{
public static int Run(string rootPath)
{
if (string.IsNullOrWhiteSpace(rootPath))
{
Console.WriteLine("UpdateBlogPaths: rootPath is required.");
return 1;
}
DataAccess.EnsureTTFileHelperColumnsExist();
string indexPath = Path.Combine(rootPath, "Index");
if (!Directory.Exists(indexPath))
{
Console.WriteLine($"Index folder not found at: {indexPath}");
return 1;
}
Console.WriteLine($"Scanning Index folder: {indexPath}");
var blogFiles = Directory.GetFiles(indexPath, "*.tumblr")
.Concat(Directory.GetFiles(indexPath, "*.tmblrpriv"))
.ToList();
Console.WriteLine($"Found {blogFiles.Count} blog metadata files");
int updatedCount = 0;
foreach (var blogFile in blogFiles)
{
try
{
string blogName = Path.GetFileNameWithoutExtension(blogFile);
string jsonContent = File.ReadAllText(blogFile);
using JsonDocument doc = JsonDocument.Parse(jsonContent);
JsonElement root = doc.RootElement;
if (root.TryGetProperty("FileDownloadLocation", out JsonElement locationElement))
{
string? fileDownloadLocation = locationElement.GetString();
if (!string.IsNullOrWhiteSpace(fileDownloadLocation))
{
DataAccess.SetBlogTTFolderPath(blogName, fileDownloadLocation);
updatedCount++;
Console.WriteLine($"Updated {blogName}: {fileDownloadLocation}");
}
}
else
{
Console.WriteLine($"No FileDownloadLocation found in {blogFile}");
}
}
catch (Exception ex)
{
Console.WriteLine($"Error processing {blogFile}: {ex.Message}");
}
}
Console.WriteLine($"\nUpdated {updatedCount} blogs with TTFolderPath");
return 0;
}
}
}
+7 -1
View File
@@ -11,7 +11,13 @@
"ContainsList": "zombaee,zomb-eh,ahzombae,thebugandme,lovingbabybug,h4rdspot", "ContainsList": "zombaee,zomb-eh,ahzombae,thebugandme,lovingbabybug,h4rdspot",
"PostIDToExclude": "730084470076686336,726309940037287936,639974102120235008,184045126218", "PostIDToExclude": "730084470076686336,726309940037287936,639974102120235008,184045126218",
"EnableFileLogging": false, "EnableFileLogging": false,
"LogTraversalRecordImports": false "LogTraversalRecordImports": false,
"LikesRefreshCooldownDays": 7,
"PathTTRoot": "",
"PathTTBackup": "",
"PathPrefixes": "prefixes.txt",
"PathCorrectionReport": "correction_report.txt",
"PathCorrectionApplied": "correction_applied.txt"
}, },
"TumblrApi": { "TumblrApi": {
"ConsumerKey": "PtsBCGumcsgihyynxUh8b47Jfmi7uXhEIOU46bmhdsUSJ5mEP3", "ConsumerKey": "PtsBCGumcsgihyynxUh8b47Jfmi7uXhEIOU46bmhdsUSJ5mEP3",
+20
View File
@@ -0,0 +1,20 @@
Post ID
reblog URL
Date
Has Image
Post URL
Slug
Reblog Key
Reblog Name
Summary
Quote
Body
Tags
Link
Photo URL
Photo Caption
Downloaded Files
Audio Caption
Question
Answer
Title
-52
View File
@@ -1,52 +0,0 @@
@echo off
REM Batch file to run URLNotesGrabberCORE 500 times in a loop
setlocal enabledelayedexpansion
REM Set the path to the executable
REM Update this path if your executable is in a different location
set APP_PATH=URLNotesGrabberCORE.exe
REM Check if the executable exists
if not exist "%APP_PATH%" (
echo Error: %APP_PATH% not found in the current directory.
echo Please ensure the executable is in the same directory as this batch file,
echo or update the APP_PATH variable with the correct path.
pause
exit /b 1
)
REM Loop counter
set ITERATIONS=500
set COUNTER=0
echo Starting to run %APP_PATH% %ITERATIONS% times...
echo.
:LOOP
set /a COUNTER+=1
echo [%COUNTER%/%ITERATIONS%] Running iteration %COUNTER%...
echo Started at: %date% %time%
REM Run the application with -replies option
call "%APP_PATH%" -replies
REM Check if the application ran successfully
if errorlevel 1 (
echo Warning: Application exited with error code !ERRORLEVEL! on iteration %COUNTER%
) else (
echo Iteration %COUNTER% completed successfully.
)
echo Completed at: %date% %time%
echo.
REM Check if we've reached 500 iterations
if %COUNTER% lss %ITERATIONS% (
goto LOOP
)
echo.
echo Completed all %ITERATIONS% iterations!
echo.
pause
+199
View File
@@ -0,0 +1,199 @@
-- ============================================================================
-- verify-db-schema.sql
--
-- Purpose: Verify that a TL.db (e.g. a restored backup) has every column the
-- current URLNotesGrabberCORE code expects. The app has NO startup
-- migration: missing columns only get added when specific modes run,
-- and a referenced-but-missing column causes a "no such column" crash.
--
-- How to use (DB Browser for SQLite):
-- 1. File > Open Database -> pick the restored backup.
-- 2. Execute SQL tab. Run SECTION 1 (it is read-only).
-- * Zero rows from every query = schema is fully aligned, you're done.
-- * Rows in "MISSING COLUMNS" = copy the run_this_to_fix text.
-- 3. If columns are missing: KEEP A COPY OF THE BACKUP FIRST, then go to
-- SECTION 2, uncomment ONLY the ALTER lines that match the report, and run.
-- 4. Re-run SECTION 1 to confirm zero rows.
--
-- This script never UPDATEs/DELETEs/DROPs. In particular it deliberately does
-- NOT replicate the likes-reset that the app's -likes migration performs
-- (DataAccess.cs:375), so existing likes high-water marks are preserved.
-- ============================================================================
-- ============================================================================
-- SECTION 1 -- VERIFICATION (read-only)
-- ============================================================================
-- Expected schema for the current code version.
-- alter_stmt is a runnable ALTER for additively-fixable columns; for base
-- columns it is a 'MANUAL REVIEW' note (a missing base column means the backup
-- predates the table's creation or is damaged -- do not blindly auto-add).
WITH expected(tbl, col, alter_stmt) AS (
VALUES
-- Posts (base columns: manual review if missing)
('Posts','BlogName', 'MANUAL REVIEW - base/PK column missing'),
('Posts','PostID', 'MANUAL REVIEW - base/PK column missing'),
('Posts','HasNotesGathered', 'MANUAL REVIEW - base column missing'),
('Posts','reblogURL', 'MANUAL REVIEW - base column missing'),
('Posts','NotFound', 'MANUAL REVIEW - base column missing'),
('Posts','PostDate', 'MANUAL REVIEW - base column missing'),
('Posts','NotesGatheredDateTime', 'MANUAL REVIEW - base column missing'),
('Posts','HasImage', 'MANUAL REVIEW - base column missing'),
('Posts','PostURL', 'MANUAL REVIEW - base column missing'),
('Posts','Slug', 'MANUAL REVIEW - base column missing'),
('Posts','ReblogKey', 'MANUAL REVIEW - base column missing'),
('Posts','ReblogName', 'MANUAL REVIEW - base column missing'),
('Posts','Summary', 'MANUAL REVIEW - base column missing'),
('Posts','Quote', 'MANUAL REVIEW - base column missing'),
('Posts','Body', 'MANUAL REVIEW - base column missing'),
('Posts','Tags', 'MANUAL REVIEW - base column missing'),
('Posts','Link', 'MANUAL REVIEW - base column missing'),
('Posts','PhotoURL', 'MANUAL REVIEW - base column missing'),
('Posts','PhotoCaption', 'MANUAL REVIEW - base column missing'),
('Posts','DownloadedFiles', 'MANUAL REVIEW - base column missing'),
('Posts','AudioCaption', 'MANUAL REVIEW - base column missing'),
('Posts','Question', 'MANUAL REVIEW - base column missing'),
('Posts','Answer', 'MANUAL REVIEW - base column missing'),
('Posts','Title', 'MANUAL REVIEW - base column missing'),
('Posts','ByLikes', 'MANUAL REVIEW - base column missing'),
('Posts','RootBlogName', 'MANUAL REVIEW - base column missing'),
('Posts','RootURL', 'MANUAL REVIEW - base column missing'),
('Posts','DateModified', 'MANUAL REVIEW - base column missing'),
('Posts','DateCreated', 'MANUAL REVIEW - base column missing'),
-- Posts (additive migration column, auto-fixable)
('Posts','PostType', 'ALTER TABLE Posts ADD COLUMN PostType TEXT;'),
-- Blogs (base columns: manual review if missing)
('Blogs','BlogName', 'MANUAL REVIEW - base/PK column missing'),
('Blogs','HasBeenOutput', 'MANUAL REVIEW - base column missing'),
('Blogs','IsActive', 'MANUAL REVIEW - base column missing'),
('Blogs','DateAdded', 'MANUAL REVIEW - base column missing'),
('Blogs','ByLikes', 'MANUAL REVIEW - base column missing'),
('Blogs','DateModified', 'MANUAL REVIEW - base column missing'),
('Blogs','DateCreated', 'MANUAL REVIEW - base column missing'),
-- Blogs (additive migration columns, auto-fixable)
('Blogs','LikesPulled', 'ALTER TABLE Blogs ADD COLUMN LikesPulled INTEGER DEFAULT 0;'),
('Blogs','LikesCursor', 'ALTER TABLE Blogs ADD COLUMN LikesCursor INTEGER DEFAULT 0;'),
('Blogs','LikesNewestTimestamp', 'ALTER TABLE Blogs ADD COLUMN LikesNewestTimestamp INTEGER DEFAULT 0;'),
('Blogs','LikesLastRefreshed', 'ALTER TABLE Blogs ADD COLUMN LikesLastRefreshed INTEGER DEFAULT 0;'),
('Blogs','LikesLastNewCount', 'ALTER TABLE Blogs ADD COLUMN LikesLastNewCount INTEGER DEFAULT 0;'),
('Blogs','TTFolderPath', 'ALTER TABLE Blogs ADD COLUMN TTFolderPath TEXT;'),
-- Notes (base columns: manual review if missing)
('Notes','RootBlogName', 'MANUAL REVIEW - base/PK column missing'),
('Notes','PostID', 'MANUAL REVIEW - base/PK column missing'),
('Notes','NoteBlogName', 'MANUAL REVIEW - base/PK column missing'),
('Notes','TimeStamp', 'MANUAL REVIEW - base/PK column missing'),
('Notes','Type', 'MANUAL REVIEW - base/PK column missing'),
('Notes','DatetimeCrawled', 'MANUAL REVIEW - base column missing'),
('Notes','DateModified', 'MANUAL REVIEW - base column missing'),
('Notes','DateCreated', 'MANUAL REVIEW - base column missing'),
-- Notes (additive migration column, auto-fixable)
('Notes','replyText', 'ALTER TABLE Notes ADD COLUMN replyText TEXT DEFAULT ''.'';'),
-- DailyAPICount (base columns)
('DailyAPICount','Date', 'MANUAL REVIEW - base/PK column missing'),
('DailyAPICount','APICount', 'MANUAL REVIEW - base column missing'),
-- ApiKeyPoolState (created at runtime by EnsureApiKeyPoolTables; auto-fixable by re-running app, but safe to add)
('ApiKeyPoolState','KeyName', 'MANUAL REVIEW - run app once to auto-create ApiKeyPool tables'),
('ApiKeyPoolState','RetryUntil', 'MANUAL REVIEW - run app once to auto-create ApiKeyPool tables'),
('ApiKeyPoolMeta','Id', 'MANUAL REVIEW - run app once to auto-create ApiKeyPool tables'),
('ApiKeyPoolMeta','LastIndex', 'MANUAL REVIEW - run app once to auto-create ApiKeyPool tables')
),
actual(tbl, col) AS (
SELECT 'Posts', name FROM pragma_table_info('Posts')
UNION ALL SELECT 'Blogs', name FROM pragma_table_info('Blogs')
UNION ALL SELECT 'Notes', name FROM pragma_table_info('Notes')
UNION ALL SELECT 'DailyAPICount', name FROM pragma_table_info('DailyAPICount')
UNION ALL SELECT 'ApiKeyPoolState', name FROM pragma_table_info('ApiKeyPoolState')
UNION ALL SELECT 'ApiKeyPoolMeta', name FROM pragma_table_info('ApiKeyPoolMeta')
)
-- 1a. MISSING COLUMNS: columns the code needs that the DB does not have.
-- Zero rows = good. Otherwise copy run_this_to_fix into SECTION 2.
SELECT
e.tbl AS table_name,
e.col AS missing_column,
e.alter_stmt AS run_this_to_fix
FROM expected e
LEFT JOIN actual a
ON a.tbl = e.tbl AND lower(a.col) = lower(e.col)
WHERE a.col IS NULL
ORDER BY (e.alter_stmt LIKE 'ALTER%') DESC, e.tbl, e.col;
-- 1b. MISSING TABLES: expected tables that don't exist at all in this DB.
-- Zero rows = good.
WITH expected_tables(tbl) AS (
VALUES ('Posts'),('Blogs'),('Notes'),('DailyAPICount'),
('ApiKeyPoolState'),('ApiKeyPoolMeta')
)
SELECT et.tbl AS missing_table
FROM expected_tables et
WHERE NOT EXISTS (
SELECT 1 FROM sqlite_master
WHERE type = 'table' AND lower(name) = lower(et.tbl)
)
ORDER BY et.tbl;
-- 1c. EXTRA / UNEXPECTED COLUMNS: present in the DB but not in the expected
-- list above. Informational only -- e.g. a NEWER backup, or a column this
-- script's expected-list hasn't been updated for. Not an error by itself.
WITH expected(tbl, col) AS (
VALUES
('Posts','BlogName'),('Posts','PostID'),('Posts','HasNotesGathered'),('Posts','reblogURL'),
('Posts','NotFound'),('Posts','PostDate'),('Posts','NotesGatheredDateTime'),('Posts','HasImage'),
('Posts','PostURL'),('Posts','Slug'),('Posts','ReblogKey'),('Posts','ReblogName'),('Posts','Summary'),
('Posts','Quote'),('Posts','Body'),('Posts','Tags'),('Posts','Link'),('Posts','PhotoURL'),
('Posts','PhotoCaption'),('Posts','DownloadedFiles'),('Posts','AudioCaption'),('Posts','Question'),
('Posts','Answer'),('Posts','Title'),('Posts','ByLikes'),('Posts','RootBlogName'),('Posts','RootURL'),
('Posts','DateModified'),('Posts','DateCreated'),('Posts','PostType'),
('Blogs','BlogName'),('Blogs','HasBeenOutput'),('Blogs','IsActive'),('Blogs','DateAdded'),
('Blogs','ByLikes'),('Blogs','DateModified'),('Blogs','DateCreated'),('Blogs','LikesPulled'),
('Blogs','LikesCursor'),('Blogs','LikesNewestTimestamp'),('Blogs','LikesLastRefreshed'),
('Blogs','LikesLastNewCount'),('Blogs','TTFolderPath'),
('Notes','RootBlogName'),('Notes','PostID'),('Notes','NoteBlogName'),('Notes','TimeStamp'),
('Notes','Type'),('Notes','DatetimeCrawled'),('Notes','DateModified'),('Notes','DateCreated'),
('Notes','replyText'),
('DailyAPICount','Date'),('DailyAPICount','APICount'),
('ApiKeyPoolState','KeyName'),('ApiKeyPoolState','RetryUntil'),
('ApiKeyPoolMeta','Id'),('ApiKeyPoolMeta','LastIndex')
),
actual(tbl, col) AS (
SELECT 'Posts', name FROM pragma_table_info('Posts')
UNION ALL SELECT 'Blogs', name FROM pragma_table_info('Blogs')
UNION ALL SELECT 'Notes', name FROM pragma_table_info('Notes')
UNION ALL SELECT 'DailyAPICount', name FROM pragma_table_info('DailyAPICount')
UNION ALL SELECT 'ApiKeyPoolState', name FROM pragma_table_info('ApiKeyPoolState')
UNION ALL SELECT 'ApiKeyPoolMeta', name FROM pragma_table_info('ApiKeyPoolMeta')
)
SELECT a.tbl AS table_name, a.col AS unexpected_column
FROM actual a
LEFT JOIN expected e
ON e.tbl = a.tbl AND lower(e.col) = lower(a.col)
WHERE e.col IS NULL
ORDER BY a.tbl, a.col;
-- ============================================================================
-- SECTION 2 -- FIX (opt-in, additive only)
--
-- Run ONLY the lines that query 1a flagged with an ALTER statement.
-- KEEP A COPY OF THE BACKUP FIRST. SQLite has no "ADD COLUMN IF NOT EXISTS",
-- so running an ALTER for a column that already exists throws a harmless
-- "duplicate column name" error and changes nothing -- just run the flagged
-- subset. These are the 8 additive migration columns and nothing else; the
-- likes high-water-mark reset is intentionally NOT included.
-- ============================================================================
-- ALTER TABLE Posts ADD COLUMN PostType TEXT;
-- ALTER TABLE Blogs ADD COLUMN LikesPulled INTEGER DEFAULT 0;
-- ALTER TABLE Blogs ADD COLUMN LikesCursor INTEGER DEFAULT 0;
-- ALTER TABLE Blogs ADD COLUMN LikesNewestTimestamp INTEGER DEFAULT 0;
-- ALTER TABLE Blogs ADD COLUMN LikesLastRefreshed INTEGER DEFAULT 0;
-- ALTER TABLE Blogs ADD COLUMN LikesLastNewCount INTEGER DEFAULT 0;
-- ALTER TABLE Blogs ADD COLUMN TTFolderPath TEXT;
-- ALTER TABLE Notes ADD COLUMN replyText TEXT DEFAULT '.';