Author SHA1 Message Date
jim c8c43c4918 Merge branch 'claude/datetime-format-consistency-b68faa' into master 2026-09-21 14:07:29 -05:00
jimandClaude Opus 5 2948a4aff0 fix(posts): normalize PostDate to yyyy-MM-dd HH:mm:ss GMT
Text-file exports can carry RFC 1123 dates ("Fri, 14 Feb 2025 15:20:09 GMT"),
which sort on the weekday name and fall outside --fromDate/--toDate
comparisons. PostDates.Normalize converts them on every PostDate write path
(AddPost, UpsertPostFromTextFile, UpdatePostContentFields).
normalize-postdate.sql fixes the 4 existing rows.

Co-Authored-By: Claude Opus 5 <[email protected]>
2026-09-21 14:07:05 -05:00
jimandClaude Sonnet 5 b8231d6a4c Merge branch 'claude/session-fa480e' into master
fix(collect): exclude posts from IsActive=0 blogs in --collect 1

Co-Authored-By: Claude Sonnet 5 <[email protected]>
2026-09-03 14:35:57 -05:00
jimandClaude Sonnet 5 bbf05b3863 fix(collect): exclude posts from IsActive=0 blogs in --collect 1
Blogs.IsActive is the crawler's work-selection flag (Rolodex removal
sets it to 0) and is independent of Posts.IsActive/Notes.IsActive --
deactivating a blog never touches its posts' own IsActive column, so
--collect 1 kept re-queuing posts for blogs that had been deactivated.

Join Blogs into the PostsWithCount CTE's source filter and require
COALESCE(BL.IsActive, 1) = 1. Since the hardcoded zomb-eh re-queue
branch reads from PostsWithCount rather than Posts directly, it now
inherits this filter automatically -- if zomb-eh is ever deactivated,
its rows disappear from PostsWithCount and the union branch
contributes nothing, with no special-case code needed.

Co-Authored-By: Claude Sonnet 5 <[email protected]>
2026-09-03 14:34:57 -05:00
jim ea2afc9d35 feat(collect): add --toDate upper bound for zomb-eh re-queue branch
Mirrors --fromDate: bounds the zomb-eh periodic re-queue branch by
PostDate <= the given date. Stacks alongside --fromDate and the
cooldown clause rather than replacing either, so --fromDate/--toDate
and --force can all be combined, and each applies with or without
--force.

Named --toDate rather than --end to pair with --fromDate -- --start
already exists as an unrelated flag (resume folder traversal at a
blog name).
2026-08-25 12:51:32 -05:00
jim 7bf270e47d feat(collect): add --fromDate lower bound for zomb-eh re-queue branch
The zomb-eh periodic re-queue branch in GetPosts had no lower bound on
the post's original PostDate -- it re-queued every already-collected
zomb-eh post past the 3-day cooldown, regardless of age.

Add --fromDate <datetime> to bound that branch by PostDate >= the given
date. It stacks with the existing cooldown clause rather than replacing
it, so it applies the same way whether or not --force also drops the
cooldown.
2026-08-24 14:42:43 -05:00
jim 58d7b1d05e Merge branch 'claude/collect-command-issue-09fcb8' into master 2026-08-23 04:34:50 -05:00
6 changed files with 161 additions and 16 deletions
+1 -1
View File
@@ -31,7 +31,7 @@ dotnet run -- --test [blogname] [postID] # Test API for specific post
- `--test [blogname] [postID]`: Test API note collection - `--test [blogname] [postID]`: Test API note collection
- `--posts`: Export post blogs to file - `--posts`: Export post blogs to file
- `--blogs`: Export blog list to file - `--blogs`: Export blog list to file
- `--collect [0|1] [datetime] [blogname]`: Collect notes for posts in DB. Optional `blogname` restricts the run to one blog (exact match), e.g. `--collect 1 zomb-eh`. Add `--force` to ignore the periodic re-collect cooldown so already-collected posts are re-queued immediately (mode 1 only) - `--collect [0|1] [datetime] [blogname]`: Collect notes for posts in DB. Optional `blogname` restricts the run to one blog (exact match), e.g. `--collect 1 zomb-eh`. Add `--force` to ignore the periodic re-collect cooldown so already-collected posts are re-queued immediately (mode 1 only). Add `--fromDate <datetime>` / `--toDate <datetime>` to only re-queue already-collected posts whose original PostDate is on/after / on/before that date (mode 1 only; either or both may be given; applies with or without `--force`)
- `--blogsR`: Export reply blogs to file - `--blogsR`: Export reply blogs to file
- `--blogsO [start] [stop]`: Export blogs within range - `--blogsO [start] [stop]`: Export blogs within range
+48 -9
View File
@@ -597,6 +597,7 @@ namespace URLNotesGrabberCORE
{ {
DBPath ??= GetDefaultDbPath(); DBPath ??= GetDefaultDbPath();
postType = PostTypes.Normalize(postType); postType = PostTypes.Normalize(postType);
postDate = PostDates.Normalize(postDate)!;
try { AddBlog(blogName, byLikes, DBPath); } catch { } try { AddBlog(blogName, byLikes, DBPath); } catch { }
try { UpdatePostSetDate(blogName, postID, postDate, DBPath); } catch { } try { UpdatePostSetDate(blogName, postID, postDate, DBPath); } catch { }
try { UpdatePost(blogName, postID, reblogURL, postDate, postURL, slug, reblogKey, reblogName, summary, quote, body, tags, link, photoURL, photoCaption, downloadedFiles, audioCaption, question, answer, title, hasImage, byLikes, DBPath, rootBlogName, rootURL, postType); } catch { } try { UpdatePost(blogName, postID, reblogURL, postDate, postURL, slug, reblogKey, reblogName, summary, quote, body, tags, link, photoURL, photoCaption, downloadedFiles, audioCaption, question, answer, title, hasImage, byLikes, DBPath, rootBlogName, rootURL, postType); } catch { }
@@ -859,9 +860,11 @@ namespace URLNotesGrabberCORE
/// </summary> /// </summary>
/// <param name="withoutNotesOnly"></param> /// <param name="withoutNotesOnly"></param>
/// <param name="ignoreRefreshCooldown">Drops the age gate on the periodic re-queue branch (--force).</param> /// <param name="ignoreRefreshCooldown">Drops the age gate on the periodic re-queue branch (--force).</param>
/// <param name="fromDate">Lower bound on the *original post's* PostDate for the periodic re-queue branch (--fromDate). Independent of ignoreRefreshCooldown -- applies whether or not --force is also given.</param>
/// <param name="toDate">Upper bound on the *original post's* PostDate for the periodic re-queue branch (--toDate). Same independence from ignoreRefreshCooldown as fromDate.</param>
/// <param name="DBPath"></param> /// <param name="DBPath"></param>
/// <returns>blogName, postID, lastNoteTimestamp, notesGatheredTimestamp</returns> /// <returns>blogName, postID, lastNoteTimestamp, notesGatheredTimestamp</returns>
public static List<Tuple<string, long, long, long>> GetPosts(bool withoutNotesOnly = false, DateTime? beforeDate = null, string? blogName = null, bool ignoreRefreshCooldown = false, string? DBPath = null) public static List<Tuple<string, long, long, long>> GetPosts(bool withoutNotesOnly = false, DateTime? beforeDate = null, string? blogName = null, bool ignoreRefreshCooldown = false, DateTime? fromDate = null, DateTime? toDate = null, string? DBPath = null)
{ {
DBPath ??= GetDefaultDbPath(); DBPath ??= GetDefaultDbPath();
using SQLiteConnection connection = new SQLiteConnection("Data Source=" + DBPath); using SQLiteConnection connection = new SQLiteConnection("Data Source=" + DBPath);
@@ -886,16 +889,25 @@ namespace URLNotesGrabberCORE
beforeDateFilter = $"WHERE (U.NotesGatheredDateTime < {unixTimestamp} OR U.NotesGatheredDateTime IS NULL)" + Environment.NewLine; beforeDateFilter = $"WHERE (U.NotesGatheredDateTime < {unixTimestamp} OR U.NotesGatheredDateTime IS NULL)" + Environment.NewLine;
} }
// Blogs.IsActive is the crawler's work-selection flag (Rolodex removal sets it to 0)
// and is independent of Posts.IsActive/Notes.IsActive -- deactivating a blog does not
// touch its posts' own IsActive column. AndIsActive("Posts", ...) above therefore does
// not catch a deactivated blog; this join against the source rows is what does, so a
// blog taken IsActive = 0 in Blogs stops being re-queued by --collect 1 even if its
// posts were never individually marked inactive. Blogs.BlogName is that table's PRIMARY
// KEY, so the join rides an index rather than scanning it.
//
// Same WHERE/AND juggling WhereIsActive does, extended to the optional blog predicate: // Same WHERE/AND juggling WhereIsActive does, extended to the optional blog predicate:
// either clause may be absent, so the first one present has to open the WHERE. // either clause may be absent, so the first one present has to open the WHERE.
string sourceClause = AndIsActive("Posts", "P", DBPath) + (filterByBlog ? " AND P.BlogName = @blogName" : string.Empty); string sourceClause = AndIsActive("Posts", "P", DBPath) + " AND COALESCE(BL.IsActive, 1) = 1" + (filterByBlog ? " AND P.BlogName = @blogName" : string.Empty);
string sourceFilter = sourceClause.Length == 0 ? string.Empty : " WHERE" + sourceClause.Substring(" AND".Length); string sourceFilter = sourceClause.Length == 0 ? string.Empty : " WHERE" + sourceClause.Substring(" AND".Length);
// The zomb-eh branch re-queues that blog's *already collected* posts every 3 days. It needs // The zomb-eh branch re-queues that blog's *already collected* posts every 3 days. It needs
// no blog-filter handling of its own: it reads PostsWithCount, which the filter has already // no blog-filter or IsActive handling of its own: it reads PostsWithCount, which the source
// scoped, so it contributes its rows when the filter names zomb-eh and nothing otherwise. // filter above -- Blogs.IsActive included -- has already scoped, so it contributes its rows
// That keeps a filtered worklist a strict subset of the unfiltered one -- "--collect 1 X" // only when zomb-eh itself is still IsActive = 1 there. That keeps a filtered worklist a
// returns exactly the rows "--collect 1" would have returned for X. // strict subset of the unfiltered one -- "--collect 1 X" returns exactly the rows
// "--collect 1" would have returned for X.
// //
// --force drops the age gate only. NotFound = 0 and the IsActive/blog scoping above still // --force drops the age gate only. NotFound = 0 and the IsActive/blog scoping above still
// apply: the flag is "re-collect early", not "collect rows every other path excludes". // apply: the flag is "re-collect early", not "collect rows every other path excludes".
@@ -903,6 +915,22 @@ namespace URLNotesGrabberCORE
? string.Empty ? string.Empty
: " AND NotesGatheredDateTime < unixepoch('now', 'localtime', '-3 days')" + Environment.NewLine; : " AND NotesGatheredDateTime < unixepoch('now', 'localtime', '-3 days')" + Environment.NewLine;
// --fromDate bounds the *original post's* PostDate, not the re-collect cooldown --
// it stacks with refreshCooldownClause instead of replacing it, so it applies the
// same way whether or not --force also dropped the cooldown. A NULL PostDate never
// satisfies ">=" and is excluded, same as an unfiltered run would still include it
// (there's nothing to compare here, so this only narrows, never widens, the result).
string fromDateClause = fromDate.HasValue
? " AND PostDate >= @fromDate" + Environment.NewLine
: string.Empty;
// --toDate is the same deal, mirrored: stacks alongside fromDateClause/
// refreshCooldownClause rather than replacing either, so --fromDate and --toDate
// can be given together (or alone) and both hold with or without --force.
string toDateClause = toDate.HasValue
? " AND PostDate <= @toDate" + Environment.NewLine
: string.Empty;
string refreshBranch = string refreshBranch =
"" + Environment.NewLine + "" + Environment.NewLine +
" UNION " + Environment.NewLine + " UNION " + Environment.NewLine +
@@ -917,7 +945,9 @@ namespace URLNotesGrabberCORE
" FROM PostsWithCount" + Environment.NewLine + " FROM PostsWithCount" + Environment.NewLine +
" WHERE BlogName = 'zomb-eh'" + Environment.NewLine + " WHERE BlogName = 'zomb-eh'" + Environment.NewLine +
" AND NotFound = 0" + Environment.NewLine + " AND NotFound = 0" + Environment.NewLine +
refreshCooldownClause; refreshCooldownClause +
fromDateClause +
toDateClause;
sql = "WITH PostsWithCount AS" + Environment.NewLine + sql = "WITH PostsWithCount AS" + Environment.NewLine +
"(" + Environment.NewLine + "(" + Environment.NewLine +
@@ -930,7 +960,7 @@ namespace URLNotesGrabberCORE
" P.HasNotesGathered," + Environment.NewLine + " P.HasNotesGathered," + Environment.NewLine +
" P.NotFound," + Environment.NewLine + " P.NotFound," + Environment.NewLine +
" P.PostDate" + Environment.NewLine + " P.PostDate" + Environment.NewLine +
" FROM Posts P" + sourceFilter + Environment.NewLine + " FROM Posts P LEFT JOIN Blogs BL ON BL.BlogName = P.BlogName" + sourceFilter + Environment.NewLine +
")," + Environment.NewLine + ")," + Environment.NewLine +
"Unioned AS" + Environment.NewLine + "Unioned AS" + Environment.NewLine +
"(" + Environment.NewLine + "(" + Environment.NewLine +
@@ -1001,6 +1031,13 @@ namespace URLNotesGrabberCORE
if (filterByBlog) if (filterByBlog)
command.Parameters.AddWithValue("@blogName", blogName); command.Parameters.AddWithValue("@blogName", blogName);
// Only ever referenced by the zomb-eh refresh branch, which only exists when
// withoutNotesOnly is true -- harmless to bind unconditionally otherwise.
if (fromDate.HasValue)
command.Parameters.AddWithValue("@fromDate", fromDate.Value.ToString("yyyy-MM-dd HH:mm:ss"));
if (toDate.HasValue)
command.Parameters.AddWithValue("@toDate", toDate.Value.ToString("yyyy-MM-dd HH:mm:ss"));
using (SQLiteDataReader reader = command.ExecuteReader()) using (SQLiteDataReader reader = command.ExecuteReader())
{ {
while (reader.Read()) while (reader.Read())
@@ -2213,6 +2250,7 @@ namespace URLNotesGrabberCORE
// column. PostType is used as an output filename, so this is the invariant that keeps // column. PostType is used as an output filename, so this is the invariant that keeps
// a stray value from becoming a stray file. // a stray value from becoming a stray file.
postType = PostTypes.Normalize(postType); postType = PostTypes.Normalize(postType);
postDate = PostDates.Normalize(postDate);
try { AddBlog(blogName, false, DBPath); } catch { } try { AddBlog(blogName, false, DBPath); } catch { }
SQLiteConnection connection; SQLiteConnection connection;
@@ -2584,10 +2622,11 @@ namespace URLNotesGrabberCORE
if (string.IsNullOrWhiteSpace(kvp.Value)) continue; if (string.IsNullOrWhiteSpace(kvp.Value)) continue;
string? column = MapPrefixToColumn(kvp.Key); string? column = MapPrefixToColumn(kvp.Key);
if (column == null) continue; if (column == null) continue;
string value = column == "PostDate" ? PostDates.Normalize(kvp.Value)! : kvp.Value;
string paramName = "@p" + parameters.Count; string paramName = "@p" + parameters.Count;
setClauses.Add($"{column} = {paramName}"); setClauses.Add($"{column} = {paramName}");
changedClauses.Add($"IFNULL({column}, '') <> {paramName}"); changedClauses.Add($"IFNULL({column}, '') <> {paramName}");
parameters.Add((paramName, kvp.Value)); parameters.Add((paramName, value));
} }
if (setClauses.Count == 0) return false; if (setClauses.Count == 0) return false;
+27
View File
@@ -0,0 +1,27 @@
using System;
using System.Globalization;
namespace URLNotesGrabberCORE
{
/// <summary>
/// The single format for Posts.PostDate: "yyyy-MM-dd HH:mm:ss GMT", which is what the
/// Tumblr API sends and what nearly every row holds. Text-file exports can carry the
/// RFC 1123 form ("Fri, 14 Feb 2025 15:20:09 GMT") instead, which as text sorts on its
/// weekday name and falls outside every --fromDate/--toDate range comparison.
///
/// Every path that writes PostDate routes through <see cref="Normalize"/>. Only the RFC 1123
/// form is rewritten; anything else, including the "." no-change sentinel, passes through.
/// </summary>
public static class PostDates
{
public static string? Normalize(string? value)
{
if (string.IsNullOrWhiteSpace(value)) return value;
string trimmed = value.Trim();
if (DateTime.TryParseExact(trimmed, "r", CultureInfo.InvariantCulture,
DateTimeStyles.AdjustToUniversal | DateTimeStyles.AssumeUniversal, out DateTime parsed))
return parsed.ToString("yyyy-MM-dd HH:mm:ss", CultureInfo.InvariantCulture) + " GMT";
return value;
}
}
}
+49 -5
View File
@@ -44,6 +44,8 @@ namespace URLNotesGrabberCORE
bool apiExplicitlySet = false; bool apiExplicitlySet = false;
string startFromBlogName = string.Empty; string startFromBlogName = string.Empty;
bool forceIgnoreCooldown = false; bool forceIgnoreCooldown = false;
DateTime? fromDate = null;
DateTime? toDate = null;
List<string> filteredArgs = new List<string>(); List<string> filteredArgs = new List<string>();
for (int i = 0; i < args.Length; i++) for (int i = 0; i < args.Length; i++)
{ {
@@ -60,6 +62,34 @@ namespace URLNotesGrabberCORE
continue; continue;
} }
if (string.Equals(args[i], "--fromDate", StringComparison.OrdinalIgnoreCase))
{
if (i + 1 < args.Length && DateTime.TryParse(args[i + 1], out DateTime parsedFromDate))
{
fromDate = parsedFromDate;
i++;
}
else
{
Console.WriteLine("--Missing or unparseable date after --fromDate. Ignoring.--");
}
continue;
}
if (string.Equals(args[i], "--toDate", StringComparison.OrdinalIgnoreCase))
{
if (i + 1 < args.Length && DateTime.TryParse(args[i + 1], out DateTime parsedToDate))
{
toDate = parsedToDate;
i++;
}
else
{
Console.WriteLine("--Missing or unparseable date after --toDate. Ignoring.--");
}
continue;
}
if (string.Equals(args[i], "--api3", StringComparison.OrdinalIgnoreCase)) if (string.Equals(args[i], "--api3", StringComparison.OrdinalIgnoreCase))
{ {
apiSectionName = "TumblrApi3"; apiSectionName = "TumblrApi3";
@@ -327,7 +357,17 @@ namespace URLNotesGrabberCORE
? "--force: ignoring the periodic re-collect cooldown; already-collected posts in scope are re-queued now" ? "--force: ignoring the periodic re-collect cooldown; already-collected posts in scope are re-queued now"
: "--force: no effect in mode 0 - a full re-check already re-collects every post"); : "--force: no effect in mode 0 - a full re-check already re-collects every post");
exitCode = CollectNotes(settings.GetValue<string>("PathOutput"), withoutNotesOnly, beforeDate, managedCollectRun, collectBlogName, forceIgnoreCooldown).GetAwaiter().GetResult(); if (fromDate.HasValue)
Console.WriteLine(withoutNotesOnly
? $"--fromDate: only re-queuing already-collected posts originally posted on/after {fromDate.Value} (applies with or without --force)"
: "--fromDate: no effect in mode 0 - it only bounds the periodic re-queue branch");
if (toDate.HasValue)
Console.WriteLine(withoutNotesOnly
? $"--toDate: only re-queuing already-collected posts originally posted on/before {toDate.Value} (applies with or without --force)"
: "--toDate: no effect in mode 0 - it only bounds the periodic re-queue branch");
exitCode = CollectNotes(settings.GetValue<string>("PathOutput"), withoutNotesOnly, beforeDate, managedCollectRun, collectBlogName, forceIgnoreCooldown, fromDate, toDate).GetAwaiter().GetResult();
break; break;
case "--blogsR": //collect notes from all posts case "--blogsR": //collect notes from all posts
@@ -459,7 +499,7 @@ namespace URLNotesGrabberCORE
Console.WriteLine("--blogs\t For each Blog in DB, write blogname to file"); Console.WriteLine("--blogs\t For each Blog in DB, write blogname to file");
Console.WriteLine("--collect [0|1] [datetime] [blogname]\t Collect Notes from API. 1=only posts without notes. 0=full re-check of all posts: a single resumable pass (interrupt & relaunch to resume; stops when complete, retrigger for a new pass). Optional datetime overrides the cutoff and runs as a one-off (bypasses resume tracking). Optional blogname restricts the run to that blog (exact, case-sensitive match) and also runs as a one-off; e.g. \"--collect 1 zomb-eh\". datetime and blogname may be given in either order - use --blog=name if a blog name would otherwise parse as a date. Add --force to ignore the periodic re-collect cooldown and re-queue already-collected posts immediately (mode 1 only)."); Console.WriteLine("--collect [0|1] [datetime] [blogname]\t Collect Notes from API. 1=only posts without notes. 0=full re-check of all posts: a single resumable pass (interrupt & relaunch to resume; stops when complete, retrigger for a new pass). Optional datetime overrides the cutoff and runs as a one-off (bypasses resume tracking). Optional blogname restricts the run to that blog (exact, case-sensitive match) and also runs as a one-off; e.g. \"--collect 1 zomb-eh\". datetime and blogname may be given in either order - use --blog=name if a blog name would otherwise parse as a date. Add --force to ignore the periodic re-collect cooldown and re-queue already-collected posts immediately (mode 1 only). Add --fromDate <datetime> / --toDate <datetime> to only re-queue already-collected posts originally posted on/after / on/before that date (mode 1 only; either or both may be given; applies with or without --force).");
Console.WriteLine("--blogsR\t For each Note that is a REPLY, write blogname to file "); Console.WriteLine("--blogsR\t For each Note that is a REPLY, write blogname to file ");
@@ -473,6 +513,10 @@ namespace URLNotesGrabberCORE
Console.WriteLine("--force\t Ignore refresh cooldowns: with --likes, refresh every fully-backfilled blog; with --collect 1, re-queue already-collected posts without waiting out their cooldown"); Console.WriteLine("--force\t Ignore refresh cooldowns: with --likes, refresh every fully-backfilled blog; with --collect 1, re-queue already-collected posts without waiting out their cooldown");
Console.WriteLine("--fromDate <datetime>\t With --collect 1, only re-queue already-collected posts originally posted on/after <datetime>. Independent of --force - applies whether or not the cooldown is also bypassed.");
Console.WriteLine("--toDate <datetime>\t With --collect 1, only re-queue already-collected posts originally posted on/before <datetime>. Independent of --force; may be combined with --fromDate for a range.");
Console.WriteLine("--urldump\t Scan all posts' text columns and extract suspected URLs to configured file"); Console.WriteLine("--urldump\t Scan all posts' text columns and extract suspected URLs to configured file");
Console.WriteLine("--api3\t Use TumblrApi3 settings from appsettings.json"); Console.WriteLine("--api3\t Use TumblrApi3 settings from appsettings.json");
@@ -1350,9 +1394,9 @@ if (shouldInsert)
// blipping on one post. Past this, skipping post-by-post would just hammer a closed door. // blipping on one post. Past this, skipping post-by-post would just hammer a closed door.
const int MaxConsecutiveTransient = 10; const int MaxConsecutiveTransient = 10;
static async Task<int> CollectNotes(string outPath, bool withoutNotesOnly = true, DateTime? beforeDate = null, bool managedRun = false, string? blogName = null, bool ignoreRefreshCooldown = false) static async Task<int> CollectNotes(string outPath, bool withoutNotesOnly = true, DateTime? beforeDate = null, bool managedRun = false, string? blogName = null, bool ignoreRefreshCooldown = false, DateTime? fromDate = null, DateTime? toDate = null)
{ {
List<Tuple<string, long, long, long>> posts = DataAccess.GetPosts(withoutNotesOnly, beforeDate, blogName, ignoreRefreshCooldown); List<Tuple<string, long, long, long>> posts = DataAccess.GetPosts(withoutNotesOnly, beforeDate, blogName, ignoreRefreshCooldown, fromDate, toDate);
if (posts.Count == 0 && !string.IsNullOrWhiteSpace(blogName)) if (posts.Count == 0 && !string.IsNullOrWhiteSpace(blogName))
{ {
@@ -1458,7 +1502,7 @@ if (shouldInsert)
} }
// Re-fetch the updated list after processing the current post // Re-fetch the updated list after processing the current post
posts = DataAccess.GetPosts(withoutNotesOnly, beforeDate, blogName, ignoreRefreshCooldown); posts = DataAccess.GetPosts(withoutNotesOnly, beforeDate, blogName, ignoreRefreshCooldown, fromDate, toDate);
} }
} }
+4
View File
@@ -154,6 +154,10 @@ Notable:
on every row, and something has since started writing it. Anything that treated it as on every row, and something has since started writing it. Anything that treated it as
permanently unset, or derived the type from post content instead, should be re-examined permanently unset, or derived the type from post content instead, should be re-examined
against the live data. Rolodex still derives it. against the live data. Rolodex still derives it.
- **`PostDate` is `yyyy-MM-dd HH:mm:ss GMT`** — UTC, as the Tumblr API sends it, unlike
the local-time `DateCreated`/`DateModified`. Text-file exports may carry RFC 1123
(`Fri, 14 Feb 2025 15:20:09 GMT`); every write path runs `PostDates.Normalize` to
convert it, and `../normalize-postdate.sql` fixed the 4 rows written before that.
- `HasImage = 1` on 14,026 rows. It records that the post *had* a picture, not that a - `HasImage = 1` on 14,026 rows. It records that the post *had* a picture, not that a
usable URL was kept, so it is not a reliable predictor that anything will render. usable URL was kept, so it is not a reliable predictor that anything will render.
- `PhotoURL` is largely unused; in practice the image markup lives inside `Body`. - `PhotoURL` is largely unused; in practice the image markup lives inside `Body`.
+31
View File
@@ -0,0 +1,31 @@
-- normalize-postdate.sql
-- Rewrites Posts.PostDate values held in RFC 1123 form ("Fri, 14 Feb 2025 15:20:09 GMT")
-- into the column's canonical "yyyy-MM-dd HH:mm:ss GMT" (the Tumblr API's own format).
--
-- As of 2026-09-21 this matched 4 rows, all zombaee, from one text-file import. As text
-- they sort on the weekday name and never satisfy --fromDate / --toDate comparisons.
-- New writes are normalized in code by PostDates.Normalize, so this is a one-off.
--
-- Only PostDate changes. DateModified is left alone: the post content did not change.
--
-- HOW TO RUN: back up TL.db, then from the URLNotesGrabberCORE project folder:
-- sqlite3 TL.db < ../normalize-postdate.sql
SELECT 'before', COUNT(*) FROM Posts WHERE PostDate LIKE '___, __ ___ ____ __:__:__ GMT';
BEGIN;
UPDATE Posts
SET PostDate = substr(PostDate, 13, 4) || '-' ||
printf('%02d', (instr('JanFebMarAprMayJunJulAugSepOctNovDec', substr(PostDate, 9, 3)) + 2) / 3) || '-' ||
substr(PostDate, 6, 2) || ' ' ||
substr(PostDate, 18)
WHERE PostDate LIKE '___, __ ___ ____ __:__:__ GMT'
AND instr('JanFebMarAprMayJunJulAugSepOctNovDec', substr(PostDate, 9, 3)) % 3 = 1;
COMMIT;
-- VERIFY: expect 0, then a single shape '9999-99-99 99:99:99 GMT' (plus any NULL/blank)
SELECT 'after', COUNT(*) FROM Posts WHERE PostDate LIKE '___, %';
SELECT CASE WHEN PostDate GLOB '[0-9][0-9][0-9][0-9]-[0-9][0-9]-[0-9][0-9] [0-9][0-9]:[0-9][0-9]:[0-9][0-9] GMT'
THEN 'yyyy-MM-dd HH:mm:ss GMT' ELSE IFNULL(PostDate, '(null)') END AS shape,
COUNT(*)
FROM Posts GROUP BY 1 ORDER BY 2 DESC;