Author SHA1 Message Date
jim c8c43c4918 Merge branch 'claude/datetime-format-consistency-b68faa' into master 2026-09-21 14:07:29 -05:00
jimandClaude Opus 5 2948a4aff0 fix(posts): normalize PostDate to yyyy-MM-dd HH:mm:ss GMT
Text-file exports can carry RFC 1123 dates ("Fri, 14 Feb 2025 15:20:09 GMT"),
which sort on the weekday name and fall outside --fromDate/--toDate
comparisons. PostDates.Normalize converts them on every PostDate write path
(AddPost, UpsertPostFromTextFile, UpdatePostContentFields).
normalize-postdate.sql fixes the 4 existing rows.

Co-Authored-By: Claude Opus 5 <[email protected]>
2026-09-21 14:07:05 -05:00
jimandClaude Sonnet 5 b8231d6a4c Merge branch 'claude/session-fa480e' into master
fix(collect): exclude posts from IsActive=0 blogs in --collect 1

Co-Authored-By: Claude Sonnet 5 <[email protected]>
2026-09-03 14:35:57 -05:00
jimandClaude Sonnet 5 bbf05b3863 fix(collect): exclude posts from IsActive=0 blogs in --collect 1
Blogs.IsActive is the crawler's work-selection flag (Rolodex removal
sets it to 0) and is independent of Posts.IsActive/Notes.IsActive --
deactivating a blog never touches its posts' own IsActive column, so
--collect 1 kept re-queuing posts for blogs that had been deactivated.

Join Blogs into the PostsWithCount CTE's source filter and require
COALESCE(BL.IsActive, 1) = 1. Since the hardcoded zomb-eh re-queue
branch reads from PostsWithCount rather than Posts directly, it now
inherits this filter automatically -- if zomb-eh is ever deactivated,
its rows disappear from PostsWithCount and the union branch
contributes nothing, with no special-case code needed.

Co-Authored-By: Claude Sonnet 5 <[email protected]>
2026-09-03 14:34:57 -05:00
jim ea2afc9d35 feat(collect): add --toDate upper bound for zomb-eh re-queue branch
Mirrors --fromDate: bounds the zomb-eh periodic re-queue branch by
PostDate <= the given date. Stacks alongside --fromDate and the
cooldown clause rather than replacing either, so --fromDate/--toDate
and --force can all be combined, and each applies with or without
--force.

Named --toDate rather than --end to pair with --fromDate -- --start
already exists as an unrelated flag (resume folder traversal at a
blog name).
2026-08-25 12:51:32 -05:00
jim 7bf270e47d feat(collect): add --fromDate lower bound for zomb-eh re-queue branch
The zomb-eh periodic re-queue branch in GetPosts had no lower bound on
the post's original PostDate -- it re-queued every already-collected
zomb-eh post past the 3-day cooldown, regardless of age.

Add --fromDate <datetime> to bound that branch by PostDate >= the given
date. It stacks with the existing cooldown clause rather than replacing
it, so it applies the same way whether or not --force also drops the
cooldown.
2026-08-24 14:42:43 -05:00
jim 58d7b1d05e Merge branch 'claude/collect-command-issue-09fcb8' into master 2026-08-23 04:34:50 -05:00
jimandClaude Opus 5 f41957fd2f feat(collect): let --force ignore the re-collect cooldown
--collect 1 <blog> returned an empty worklist whenever the periodic
re-queue branch was still inside its 3-day window, with no way to ask
for the re-collect early. GetPosts now composes that age predicate
conditionally, and the existing global --force flag - already "ignore
cooldown" for --likes - drives it.

Only the age gate drops: NotFound = 0, the IsActive filter and the blog
scoping still apply. The flag is a no-op for mode 0, which re-collects
every post regardless, and says so rather than pretending to act.

Co-Authored-By: Claude Opus 5 <[email protected]>
2026-08-23 04:34:45 -05:00
6 changed files with 177 additions and 17 deletions
+1 -1
View File
@@ -31,7 +31,7 @@ dotnet run -- --test [blogname] [postID] # Test API for specific post
- `--test [blogname] [postID]`: Test API note collection
- `--posts`: Export post blogs to file
- `--blogs`: Export blog list to file
- `--collect [0|1] [datetime] [blogname]`: Collect notes for posts in DB. Optional `blogname` restricts the run to one blog (exact match), e.g. `--collect 1 zomb-eh`
- `--collect [0|1] [datetime] [blogname]`: Collect notes for posts in DB. Optional `blogname` restricts the run to one blog (exact match), e.g. `--collect 1 zomb-eh`. Add `--force` to ignore the periodic re-collect cooldown so already-collected posts are re-queued immediately (mode 1 only). Add `--fromDate <datetime>` / `--toDate <datetime>` to only re-queue already-collected posts whose original PostDate is on/after / on/before that date (mode 1 only; either or both may be given; applies with or without `--force`)
- `--blogsR`: Export reply blogs to file
- `--blogsO [start] [stop]`: Export blogs within range
+57 -10
View File
@@ -597,6 +597,7 @@ namespace URLNotesGrabberCORE
{
DBPath ??= GetDefaultDbPath();
postType = PostTypes.Normalize(postType);
postDate = PostDates.Normalize(postDate)!;
try { AddBlog(blogName, byLikes, DBPath); } catch { }
try { UpdatePostSetDate(blogName, postID, postDate, DBPath); } catch { }
try { UpdatePost(blogName, postID, reblogURL, postDate, postURL, slug, reblogKey, reblogName, summary, quote, body, tags, link, photoURL, photoCaption, downloadedFiles, audioCaption, question, answer, title, hasImage, byLikes, DBPath, rootBlogName, rootURL, postType); } catch { }
@@ -855,12 +856,15 @@ namespace URLNotesGrabberCORE
#region Gets
/// <summary>
///
///
/// </summary>
/// <param name="withoutNotesOnly"></param>
/// <param name="ignoreRefreshCooldown">Drops the age gate on the periodic re-queue branch (--force).</param>
/// <param name="fromDate">Lower bound on the *original post's* PostDate for the periodic re-queue branch (--fromDate). Independent of ignoreRefreshCooldown -- applies whether or not --force is also given.</param>
/// <param name="toDate">Upper bound on the *original post's* PostDate for the periodic re-queue branch (--toDate). Same independence from ignoreRefreshCooldown as fromDate.</param>
/// <param name="DBPath"></param>
/// <returns>blogName, postID, lastNoteTimestamp, notesGatheredTimestamp</returns>
public static List<Tuple<string, long, long, long>> GetPosts(bool withoutNotesOnly = false, DateTime? beforeDate = null, string? blogName = null, string? DBPath = null)
public static List<Tuple<string, long, long, long>> GetPosts(bool withoutNotesOnly = false, DateTime? beforeDate = null, string? blogName = null, bool ignoreRefreshCooldown = false, DateTime? fromDate = null, DateTime? toDate = null, string? DBPath = null)
{
DBPath ??= GetDefaultDbPath();
using SQLiteConnection connection = new SQLiteConnection("Data Source=" + DBPath);
@@ -885,16 +889,48 @@ namespace URLNotesGrabberCORE
beforeDateFilter = $"WHERE (U.NotesGatheredDateTime < {unixTimestamp} OR U.NotesGatheredDateTime IS NULL)" + Environment.NewLine;
}
// Blogs.IsActive is the crawler's work-selection flag (Rolodex removal sets it to 0)
// and is independent of Posts.IsActive/Notes.IsActive -- deactivating a blog does not
// touch its posts' own IsActive column. AndIsActive("Posts", ...) above therefore does
// not catch a deactivated blog; this join against the source rows is what does, so a
// blog taken IsActive = 0 in Blogs stops being re-queued by --collect 1 even if its
// posts were never individually marked inactive. Blogs.BlogName is that table's PRIMARY
// KEY, so the join rides an index rather than scanning it.
//
// Same WHERE/AND juggling WhereIsActive does, extended to the optional blog predicate:
// either clause may be absent, so the first one present has to open the WHERE.
string sourceClause = AndIsActive("Posts", "P", DBPath) + (filterByBlog ? " AND P.BlogName = @blogName" : string.Empty);
string sourceClause = AndIsActive("Posts", "P", DBPath) + " AND COALESCE(BL.IsActive, 1) = 1" + (filterByBlog ? " AND P.BlogName = @blogName" : string.Empty);
string sourceFilter = sourceClause.Length == 0 ? string.Empty : " WHERE" + sourceClause.Substring(" AND".Length);
// The zomb-eh branch re-queues that blog's *already collected* posts every 3 days. It needs
// no blog-filter handling of its own: it reads PostsWithCount, which the filter has already
// scoped, so it contributes its rows when the filter names zomb-eh and nothing otherwise.
// That keeps a filtered worklist a strict subset of the unfiltered one -- "--collect 1 X"
// returns exactly the rows "--collect 1" would have returned for X.
// no blog-filter or IsActive handling of its own: it reads PostsWithCount, which the source
// filter above -- Blogs.IsActive included -- has already scoped, so it contributes its rows
// only when zomb-eh itself is still IsActive = 1 there. That keeps a filtered worklist a
// strict subset of the unfiltered one -- "--collect 1 X" returns exactly the rows
// "--collect 1" would have returned for X.
//
// --force drops the age gate only. NotFound = 0 and the IsActive/blog scoping above still
// apply: the flag is "re-collect early", not "collect rows every other path excludes".
string refreshCooldownClause = ignoreRefreshCooldown
? string.Empty
: " AND NotesGatheredDateTime < unixepoch('now', 'localtime', '-3 days')" + Environment.NewLine;
// --fromDate bounds the *original post's* PostDate, not the re-collect cooldown --
// it stacks with refreshCooldownClause instead of replacing it, so it applies the
// same way whether or not --force also dropped the cooldown. A NULL PostDate never
// satisfies ">=" and is excluded, same as an unfiltered run would still include it
// (there's nothing to compare here, so this only narrows, never widens, the result).
string fromDateClause = fromDate.HasValue
? " AND PostDate >= @fromDate" + Environment.NewLine
: string.Empty;
// --toDate is the same deal, mirrored: stacks alongside fromDateClause/
// refreshCooldownClause rather than replacing either, so --fromDate and --toDate
// can be given together (or alone) and both hold with or without --force.
string toDateClause = toDate.HasValue
? " AND PostDate <= @toDate" + Environment.NewLine
: string.Empty;
string refreshBranch =
"" + Environment.NewLine +
" UNION " + Environment.NewLine +
@@ -909,7 +945,9 @@ namespace URLNotesGrabberCORE
" FROM PostsWithCount" + Environment.NewLine +
" WHERE BlogName = 'zomb-eh'" + Environment.NewLine +
" AND NotFound = 0" + Environment.NewLine +
" AND NotesGatheredDateTime < unixepoch('now', 'localtime', '-3 days')" + Environment.NewLine;
refreshCooldownClause +
fromDateClause +
toDateClause;
sql = "WITH PostsWithCount AS" + Environment.NewLine +
"(" + Environment.NewLine +
@@ -922,7 +960,7 @@ namespace URLNotesGrabberCORE
" P.HasNotesGathered," + Environment.NewLine +
" P.NotFound," + Environment.NewLine +
" P.PostDate" + Environment.NewLine +
" FROM Posts P" + sourceFilter + Environment.NewLine +
" FROM Posts P LEFT JOIN Blogs BL ON BL.BlogName = P.BlogName" + sourceFilter + Environment.NewLine +
")," + Environment.NewLine +
"Unioned AS" + Environment.NewLine +
"(" + Environment.NewLine +
@@ -993,6 +1031,13 @@ namespace URLNotesGrabberCORE
if (filterByBlog)
command.Parameters.AddWithValue("@blogName", blogName);
// Only ever referenced by the zomb-eh refresh branch, which only exists when
// withoutNotesOnly is true -- harmless to bind unconditionally otherwise.
if (fromDate.HasValue)
command.Parameters.AddWithValue("@fromDate", fromDate.Value.ToString("yyyy-MM-dd HH:mm:ss"));
if (toDate.HasValue)
command.Parameters.AddWithValue("@toDate", toDate.Value.ToString("yyyy-MM-dd HH:mm:ss"));
using (SQLiteDataReader reader = command.ExecuteReader())
{
while (reader.Read())
@@ -2205,6 +2250,7 @@ namespace URLNotesGrabberCORE
// column. PostType is used as an output filename, so this is the invariant that keeps
// a stray value from becoming a stray file.
postType = PostTypes.Normalize(postType);
postDate = PostDates.Normalize(postDate);
try { AddBlog(blogName, false, DBPath); } catch { }
SQLiteConnection connection;
@@ -2576,10 +2622,11 @@ namespace URLNotesGrabberCORE
if (string.IsNullOrWhiteSpace(kvp.Value)) continue;
string? column = MapPrefixToColumn(kvp.Key);
if (column == null) continue;
string value = column == "PostDate" ? PostDates.Normalize(kvp.Value)! : kvp.Value;
string paramName = "@p" + parameters.Count;
setClauses.Add($"{column} = {paramName}");
changedClauses.Add($"IFNULL({column}, '') <> {paramName}");
parameters.Add((paramName, kvp.Value));
parameters.Add((paramName, value));
}
if (setClauses.Count == 0) return false;
+27
View File
@@ -0,0 +1,27 @@
using System;
using System.Globalization;
namespace URLNotesGrabberCORE
{
/// <summary>
/// The single format for Posts.PostDate: "yyyy-MM-dd HH:mm:ss GMT", which is what the
/// Tumblr API sends and what nearly every row holds. Text-file exports can carry the
/// RFC 1123 form ("Fri, 14 Feb 2025 15:20:09 GMT") instead, which as text sorts on its
/// weekday name and falls outside every --fromDate/--toDate range comparison.
///
/// Every path that writes PostDate routes through <see cref="Normalize"/>. Only the RFC 1123
/// form is rewritten; anything else, including the "." no-change sentinel, passes through.
/// </summary>
public static class PostDates
{
public static string? Normalize(string? value)
{
if (string.IsNullOrWhiteSpace(value)) return value;
string trimmed = value.Trim();
if (DateTime.TryParseExact(trimmed, "r", CultureInfo.InvariantCulture,
DateTimeStyles.AdjustToUniversal | DateTimeStyles.AssumeUniversal, out DateTime parsed))
return parsed.ToString("yyyy-MM-dd HH:mm:ss", CultureInfo.InvariantCulture) + " GMT";
return value;
}
}
}
+57 -6
View File
@@ -44,6 +44,8 @@ namespace URLNotesGrabberCORE
bool apiExplicitlySet = false;
string startFromBlogName = string.Empty;
bool forceIgnoreCooldown = false;
DateTime? fromDate = null;
DateTime? toDate = null;
List<string> filteredArgs = new List<string>();
for (int i = 0; i < args.Length; i++)
{
@@ -60,6 +62,34 @@ namespace URLNotesGrabberCORE
continue;
}
if (string.Equals(args[i], "--fromDate", StringComparison.OrdinalIgnoreCase))
{
if (i + 1 < args.Length && DateTime.TryParse(args[i + 1], out DateTime parsedFromDate))
{
fromDate = parsedFromDate;
i++;
}
else
{
Console.WriteLine("--Missing or unparseable date after --fromDate. Ignoring.--");
}
continue;
}
if (string.Equals(args[i], "--toDate", StringComparison.OrdinalIgnoreCase))
{
if (i + 1 < args.Length && DateTime.TryParse(args[i + 1], out DateTime parsedToDate))
{
toDate = parsedToDate;
i++;
}
else
{
Console.WriteLine("--Missing or unparseable date after --toDate. Ignoring.--");
}
continue;
}
if (string.Equals(args[i], "--api3", StringComparison.OrdinalIgnoreCase))
{
apiSectionName = "TumblrApi3";
@@ -322,7 +352,22 @@ namespace URLNotesGrabberCORE
managedCollectRun = true;
}
exitCode = CollectNotes(settings.GetValue<string>("PathOutput"), withoutNotesOnly, beforeDate, managedCollectRun, collectBlogName).GetAwaiter().GetResult();
if (forceIgnoreCooldown)
Console.WriteLine(withoutNotesOnly
? "--force: ignoring the periodic re-collect cooldown; already-collected posts in scope are re-queued now"
: "--force: no effect in mode 0 - a full re-check already re-collects every post");
if (fromDate.HasValue)
Console.WriteLine(withoutNotesOnly
? $"--fromDate: only re-queuing already-collected posts originally posted on/after {fromDate.Value} (applies with or without --force)"
: "--fromDate: no effect in mode 0 - it only bounds the periodic re-queue branch");
if (toDate.HasValue)
Console.WriteLine(withoutNotesOnly
? $"--toDate: only re-queuing already-collected posts originally posted on/before {toDate.Value} (applies with or without --force)"
: "--toDate: no effect in mode 0 - it only bounds the periodic re-queue branch");
exitCode = CollectNotes(settings.GetValue<string>("PathOutput"), withoutNotesOnly, beforeDate, managedCollectRun, collectBlogName, forceIgnoreCooldown, fromDate, toDate).GetAwaiter().GetResult();
break;
case "--blogsR": //collect notes from all posts
@@ -454,7 +499,7 @@ namespace URLNotesGrabberCORE
Console.WriteLine("--blogs\t For each Blog in DB, write blogname to file");
Console.WriteLine("--collect [0|1] [datetime] [blogname]\t Collect Notes from API. 1=only posts without notes. 0=full re-check of all posts: a single resumable pass (interrupt & relaunch to resume; stops when complete, retrigger for a new pass). Optional datetime overrides the cutoff and runs as a one-off (bypasses resume tracking). Optional blogname restricts the run to that blog (exact, case-sensitive match) and also runs as a one-off; e.g. \"--collect 1 zomb-eh\". datetime and blogname may be given in either order - use --blog=name if a blog name would otherwise parse as a date.");
Console.WriteLine("--collect [0|1] [datetime] [blogname]\t Collect Notes from API. 1=only posts without notes. 0=full re-check of all posts: a single resumable pass (interrupt & relaunch to resume; stops when complete, retrigger for a new pass). Optional datetime overrides the cutoff and runs as a one-off (bypasses resume tracking). Optional blogname restricts the run to that blog (exact, case-sensitive match) and also runs as a one-off; e.g. \"--collect 1 zomb-eh\". datetime and blogname may be given in either order - use --blog=name if a blog name would otherwise parse as a date. Add --force to ignore the periodic re-collect cooldown and re-queue already-collected posts immediately (mode 1 only). Add --fromDate <datetime> / --toDate <datetime> to only re-queue already-collected posts originally posted on/after / on/before that date (mode 1 only; either or both may be given; applies with or without --force).");
Console.WriteLine("--blogsR\t For each Note that is a REPLY, write blogname to file ");
@@ -466,7 +511,11 @@ namespace URLNotesGrabberCORE
Console.WriteLine("--likes\t Fetch likes: initial backfill for new blogs, incremental refresh for blogs past cooldown. Optional blog name forces single-blog run.");
Console.WriteLine("--force\t (with --likes) Ignore cooldown and refresh every fully-backfilled blog");
Console.WriteLine("--force\t Ignore refresh cooldowns: with --likes, refresh every fully-backfilled blog; with --collect 1, re-queue already-collected posts without waiting out their cooldown");
Console.WriteLine("--fromDate <datetime>\t With --collect 1, only re-queue already-collected posts originally posted on/after <datetime>. Independent of --force - applies whether or not the cooldown is also bypassed.");
Console.WriteLine("--toDate <datetime>\t With --collect 1, only re-queue already-collected posts originally posted on/before <datetime>. Independent of --force; may be combined with --fromDate for a range.");
Console.WriteLine("--urldump\t Scan all posts' text columns and extract suspected URLs to configured file");
@@ -1345,15 +1394,17 @@ if (shouldInsert)
// blipping on one post. Past this, skipping post-by-post would just hammer a closed door.
const int MaxConsecutiveTransient = 10;
static async Task<int> CollectNotes(string outPath, bool withoutNotesOnly = true, DateTime? beforeDate = null, bool managedRun = false, string? blogName = null)
static async Task<int> CollectNotes(string outPath, bool withoutNotesOnly = true, DateTime? beforeDate = null, bool managedRun = false, string? blogName = null, bool ignoreRefreshCooldown = false, DateTime? fromDate = null, DateTime? toDate = null)
{
List<Tuple<string, long, long, long>> posts = DataAccess.GetPosts(withoutNotesOnly, beforeDate, blogName);
List<Tuple<string, long, long, long>> posts = DataAccess.GetPosts(withoutNotesOnly, beforeDate, blogName, ignoreRefreshCooldown, fromDate, toDate);
if (posts.Count == 0 && !string.IsNullOrWhiteSpace(blogName))
{
// BlogName is matched exactly, so a typo or a case mismatch looks identical to "nothing
// left to collect". Say so rather than reporting a silent, instant success.
Console.WriteLine($"No posts to collect for blog '{blogName}'. Either it is fully collected, or the name does not match a stored blog (the match is case-sensitive).");
if (withoutNotesOnly && !ignoreRefreshCooldown)
Console.WriteLine("Already-collected posts are re-queued only once their cooldown elapses; add --force to re-collect them now.");
return 0;
}
@@ -1451,7 +1502,7 @@ if (shouldInsert)
}
// Re-fetch the updated list after processing the current post
posts = DataAccess.GetPosts(withoutNotesOnly, beforeDate, blogName);
posts = DataAccess.GetPosts(withoutNotesOnly, beforeDate, blogName, ignoreRefreshCooldown, fromDate, toDate);
}
}
+4
View File
@@ -154,6 +154,10 @@ Notable:
on every row, and something has since started writing it. Anything that treated it as
permanently unset, or derived the type from post content instead, should be re-examined
against the live data. Rolodex still derives it.
- **`PostDate` is `yyyy-MM-dd HH:mm:ss GMT`** — UTC, as the Tumblr API sends it, unlike
the local-time `DateCreated`/`DateModified`. Text-file exports may carry RFC 1123
(`Fri, 14 Feb 2025 15:20:09 GMT`); every write path runs `PostDates.Normalize` to
convert it, and `../normalize-postdate.sql` fixed the 4 rows written before that.
- `HasImage = 1` on 14,026 rows. It records that the post *had* a picture, not that a
usable URL was kept, so it is not a reliable predictor that anything will render.
- `PhotoURL` is largely unused; in practice the image markup lives inside `Body`.
+31
View File
@@ -0,0 +1,31 @@
-- normalize-postdate.sql
-- Rewrites Posts.PostDate values held in RFC 1123 form ("Fri, 14 Feb 2025 15:20:09 GMT")
-- into the column's canonical "yyyy-MM-dd HH:mm:ss GMT" (the Tumblr API's own format).
--
-- As of 2026-09-21 this matched 4 rows, all zombaee, from one text-file import. As text
-- they sort on the weekday name and never satisfy --fromDate / --toDate comparisons.
-- New writes are normalized in code by PostDates.Normalize, so this is a one-off.
--
-- Only PostDate changes. DateModified is left alone: the post content did not change.
--
-- HOW TO RUN: back up TL.db, then from the URLNotesGrabberCORE project folder:
-- sqlite3 TL.db < ../normalize-postdate.sql
SELECT 'before', COUNT(*) FROM Posts WHERE PostDate LIKE '___, __ ___ ____ __:__:__ GMT';
BEGIN;
UPDATE Posts
SET PostDate = substr(PostDate, 13, 4) || '-' ||
printf('%02d', (instr('JanFebMarAprMayJunJulAugSepOctNovDec', substr(PostDate, 9, 3)) + 2) / 3) || '-' ||
substr(PostDate, 6, 2) || ' ' ||
substr(PostDate, 18)
WHERE PostDate LIKE '___, __ ___ ____ __:__:__ GMT'
AND instr('JanFebMarAprMayJunJulAugSepOctNovDec', substr(PostDate, 9, 3)) % 3 = 1;
COMMIT;
-- VERIFY: expect 0, then a single shape '9999-99-99 99:99:99 GMT' (plus any NULL/blank)
SELECT 'after', COUNT(*) FROM Posts WHERE PostDate LIKE '___, %';
SELECT CASE WHEN PostDate GLOB '[0-9][0-9][0-9][0-9]-[0-9][0-9]-[0-9][0-9] [0-9][0-9]:[0-9][0-9]:[0-9][0-9] GMT'
THEN 'yyyy-MM-dd HH:mm:ss GMT' ELSE IFNULL(PostDate, '(null)') END AS shape,
COUNT(*)
FROM Posts GROUP BY 1 ORDER BY 2 DESC;