fix: strip only trailing numeric suffix from blog folder names

NormalizeBlogFolderName removed "_1".."_9" as unanchored substrings, so a
folder suffixed past a single digit lost the wrong characters: "_10" hit
the "_1" rule and left the trailing "0" welded to the name, importing
zomb-eh_10 as blog "zomb-eh0". That name does not exist on Tumblr, so
every post imported under it 404s on --collect forever.

Anchor the strip to a trailing _<digits> instead. This also fixes blogs
whose real name contains "_1" (some_1blog no longer becomes someblog) and
folders suffixed "_0", which were not stripped at all.

Verified against the live folder tree: zomb-eh_10 is the only existing
folder whose normalized name changes.

Co-Authored-By: Claude Opus 4.8 <[email protected]>
This commit is contained in:
jim
2026-07-22 12:58:35 -05:00
co-authored by Claude Opus 4.8
parent f9e1d2100b
commit a73b597381
+5 -10
View File
@@ -582,16 +582,11 @@ namespace URLNotesGrabberCORE
protected static string NormalizeBlogFolderName(string folderName) protected static string NormalizeBlogFolderName(string folderName)
{ {
return folderName // Archive tools suffix duplicate blog folders with _1, _2, ... _10 and beyond. Strip only a
.Replace("_1", "") // trailing numeric suffix: unanchored substring removal ate the "_1" inside "_10" and left the
.Replace("_2", "") // "0" welded to the name (zomb-eh_10 -> zomb-eh0), and mangled any blog whose real name
.Replace("_3", "") // contains "_1". A blog name is never a prefix of itself plus "_<digits>", so this is safe.
.Replace("_4", "") return System.Text.RegularExpressions.Regex.Replace(folderName, @"_\d+$", "");
.Replace("_5", "")
.Replace("_6", "")
.Replace("_7", "")
.Replace("_8", "")
.Replace("_9", "");
} }
static async Task FetchAndStoreReplyText(string blogName, long postID, long timestamp) static async Task FetchAndStoreReplyText(string blogName, long postID, long timestamp)