Compare commits
12
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
8f4177a0c9 | ||
|
|
721224bc13 | ||
|
|
ef6629d86a | ||
|
|
6320e2c0c9 | ||
|
|
6136901cc7 | ||
|
|
83e35a2323 | ||
|
|
f0ccac6503 | ||
|
|
e1d2eb48c2 | ||
|
|
3c85a05afc | ||
|
|
60912c882d | ||
|
|
8a4ab2402d | ||
|
|
05ec465f74 |
@@ -47,6 +47,113 @@ say nothing about the item being fetched, so they must not be recorded as per-it
|
||||
- Long-running commands return exit 3 when a pass ends incomplete (rate-limit pause, breaker trip, or
|
||||
skipped items), so a caller can distinguish that from a clean run
|
||||
|
||||
### `IsActive` Is Not Ours To Write
|
||||
`Blogs.IsActive`, `Posts.IsActive` and `Notes.IsActive` are removal flags set by other tools
|
||||
(Rolodex). `0` means removed; anything else, including `NULL`, means live. Full detail in
|
||||
`URLNotesGrabberCORE/TL.db.md`.
|
||||
|
||||
- **Never write any `IsActive` column.** Not in an `INSERT` column list, not in an `UPDATE`,
|
||||
and never via `INSERT OR REPLACE` on these tables — that resets the column default and
|
||||
un-removes the row. Re-crawling a removed row must refresh its content and leave the flag
|
||||
where it was
|
||||
- **Filter at selection, not at write.** Every query that *selects* posts, notes or blogs
|
||||
excludes removed rows. Update statements stay keyed on a row the caller already selected;
|
||||
filtering them would spend API quota and then fail to persist the result
|
||||
- `Posts.IsActive` and `Notes.IsActive` are **optional** — they are absent from databases
|
||||
that predate them, and naming a missing column is a hard SQLite error. Compose the filter
|
||||
with `AndIsActive`/`WhereIsActive` in `DataAccess.cs`, which return
|
||||
`COALESCE(IsActive, 1) = 1` only when `HasIsActiveColumn` finds the column. `Blogs.IsActive`
|
||||
is not optional and is filtered directly
|
||||
- Do not add these columns from this app, and do not add them to the missing-column list in
|
||||
`verify-db-schema.sql`
|
||||
|
||||
### `DateModified` Tracks Real Changes Only
|
||||
`Blogs.DateModified`, `Posts.DateModified` and `Notes.DateModified` must move only when a
|
||||
column beside `DateModified` itself actually changed. Re-crawling or re-ingesting identical
|
||||
content has to leave the row — and its timestamp — untouched, or downstream consumers cannot
|
||||
tell a refreshed row from a rewritten one.
|
||||
|
||||
- Enforce it in the `WHERE` clause, not in C#. Every `UPDATE` that sets `DateModified` ends
|
||||
with an `AND (<col> <> @param OR ...)` term covering every column in its `SET` list, so
|
||||
SQLite matches zero rows on a no-op and never writes
|
||||
- Compare NULL-safely: `IFNULL(col, '') <> IFNULL(@param, '')` for text,
|
||||
`IFNULL(col, 0) <> @param` for integer flags. A bare `col <> @param` is NULL on a NULL
|
||||
column and silently skips the row that most needs writing
|
||||
- Where NULL is not equivalent to the default, spell it out. The `HasBeenOutput = 0` stamps
|
||||
use `(HasBeenOutput IS NULL OR HasBeenOutput <> 0)` because the selection queries test
|
||||
`HasBeenOutput = 0`, which a NULL would never match
|
||||
- Dynamic `SET` lists (`UpdatePostContentFields`) build the guard alongside the assignments
|
||||
so the two lists cannot drift apart
|
||||
- These statements now return 0 rows for "found but unchanged" as well as "not found".
|
||||
Callers that read `ExecuteNonQuery()` must not treat 0 as "row missing"
|
||||
|
||||
**`Posts.NotesGatheredDateTime` is crawl bookkeeping, not content.** It moves on every
|
||||
`-collect` pass and says nothing about the post, so it must never move `DateModified` on its
|
||||
own. `UpdatePostMarkNotesCollected` still writes it every pass but wraps the timestamp in
|
||||
`DateModified = CASE WHEN IFNULL(HasNotesGathered, 0) <> 1 THEN @dateModified ELSE
|
||||
DateModified END` — SQLite evaluates `SET` expressions against the pre-`UPDATE` row, so only
|
||||
the flag flipping counts as a modification. Use this shape for any column that has to be
|
||||
refreshed unconditionally without being a change. `Blogs.LikesLastRefreshed` is the
|
||||
deliberate exception: a refresh pass is treated as a real event on the blog row.
|
||||
|
||||
**`Blogs.DateAdded` is write-once.** `AddBlog`'s `INSERT` is the only place that sets it. A
|
||||
new post arriving for a known blog reopens `HasBeenOutput` but must leave `DateAdded` alone —
|
||||
a new post is not a new blog, and rewriting the column both destroys the registration date
|
||||
and makes every insert look like a change.
|
||||
|
||||
**`"."` in a `Posts` content field means "not supplied", not "empty".** `ReblogRecord`
|
||||
(`TraverseDirectory`'s parser for the local `.txt` export tree) and the `--likes` API path
|
||||
both default every content field to the literal string `"."` when their source has no value
|
||||
for it, then pass that straight to `UpdatePost`. A blog with two export folders in different
|
||||
field formats (a duplicate `_2` folder, or a Tumblr export whose field set changed over time)
|
||||
sends one record with a real `Title`/`Tags`/`Slug` and another with those fields `"."`
|
||||
because that format never had a line for them — and without a guard, re-importing both on
|
||||
every run flips the row back and forth forever, bumping `DateModified` on every pass even
|
||||
though the true content never changes.
|
||||
|
||||
- Every content column in `UpdatePost`'s `SET` list is guarded the same way `RootBlogName`/
|
||||
`RootURL` already were: `col = CASE WHEN @col = '.' THEN col ELSE @col END`. A `"."`
|
||||
parameter leaves the existing value alone instead of overwriting it
|
||||
- The change-detection `WHERE` clause carries the same exception —
|
||||
`(@col <> '.' AND IFNULL(col, '') <> @col) OR ...` — so a `"."`-only difference does not
|
||||
make the statement fire at all, and `DateModified` stays put
|
||||
- Deliberately narrow: only the literal `"."` is the sentinel. An explicit empty string from
|
||||
a real record still overwrites, same as before this fix. `postID`, `BlogName`, `hasImage`,
|
||||
`ByLikes` are not part of this convention and are unaffected
|
||||
- If a new content field is added to `Posts`/`UpdatePost`, decide explicitly whether its
|
||||
source can legitimately supply `"."` as "field absent" before deciding whether it needs
|
||||
the same `CASE` treatment — don't assume every column needs it
|
||||
|
||||
**`--ingest` (`UpsertPostFromTextFile`) uses `NULL`, not `"."`, for the same "field absent"
|
||||
convention, and reconciling exactly this kind of duplicate IS the feature's job.**
|
||||
`IngestMode` strips a trailing `_N` from the folder name before it ever reaches
|
||||
`UpsertPostFromTextFile`, so a duplicate export folder collapses onto the same `BlogName` on
|
||||
purpose — the whole point is to merge multiple differently-formatted files for the same post
|
||||
into one row. `IngestMode.G(key)` returns `null` (not `"."`) when a field's line is absent
|
||||
from a given file, `LegacyPostsDbImporter` passes `null` straight from a `NULL` source column,
|
||||
and files are walked in raw filesystem enumeration order — never sorted — so which file's call
|
||||
lands last for a given `(BlogName, PostID)` is arbitrary.
|
||||
|
||||
- Before the fix, the `UPDATE` branch set every column unconditionally, so whichever file
|
||||
processed last for a `PostID` would null out every field its own record didn't carry —
|
||||
silently erasing real `Title`/`Slug`/`Tags`/… another file had, the opposite of what
|
||||
`--ingest` exists to do. This is worse than the `"."` case above: that one only caused
|
||||
churn (the two writes canceled out); this one loses data, and which posts lose which
|
||||
fields depends on filesystem enumeration order
|
||||
- Same shape of fix, `NULL` instead of `"."` as the sentinel: `col = CASE WHEN @col IS NULL
|
||||
THEN col ELSE @col END` in the `SET` list, `(@col IS NOT NULL AND IFNULL(col, '') <> @col)
|
||||
OR ...` in the change-detection
|
||||
- Same narrow rule: only `NULL` (the field's line was never present in this file) is the
|
||||
sentinel. `G()` already distinguishes this from "present but blank" — a dictionary miss is
|
||||
`null`, an empty value after the prefix is `""` — so an explicitly blank field still
|
||||
overwrites
|
||||
- `HasImage` is **not** guarded and remains a known gap: `IngestMode` always computes a
|
||||
concrete `bool` (defaulting `false` when a file has no `Has Image:` line), so there is no
|
||||
way for this function to tell "this format says no image" from "this format doesn't report
|
||||
it at all" without changing the parameter to `bool?` and threading that through
|
||||
`IngestMode`/`LegacyPostsDbImporter`. Fix this the same way if `--ingest` is observed
|
||||
downgrading a post's `HasImage` from `1` to `0`
|
||||
|
||||
### Testing
|
||||
- No existing test suite; use xUnit if adding tests
|
||||
- Test critical logic: `ApiKeyPool` init, color parsing, config persistence
|
||||
|
||||
@@ -80,6 +80,10 @@ namespace URLNotesGrabberCORE
|
||||
private static SQLiteConnection? _importConnection;
|
||||
private static HashSet<string>? _importBlogCache;
|
||||
private static readonly object _importSessionLock = new object();
|
||||
// AddAPICount/UpdateAPICount run once per API call. A schema-level failure there repeats
|
||||
// identically every time, so log each distinct message once instead of per call.
|
||||
private static readonly HashSet<string> _apiCountFailuresLogged = new HashSet<string>();
|
||||
private static readonly object _apiCountFailureLock = new object();
|
||||
|
||||
static DataAccess()
|
||||
{
|
||||
@@ -102,6 +106,10 @@ namespace URLNotesGrabberCORE
|
||||
}
|
||||
}
|
||||
|
||||
// The database every DataAccess call defaults to, exposed so modes can report
|
||||
// which file they actually read when their results are surprising.
|
||||
public static string GetActiveDbPath() => GetDefaultDbPath();
|
||||
|
||||
private static string GetDefaultDbPath()
|
||||
{
|
||||
if (_cachedDbPath != null)
|
||||
@@ -120,6 +128,91 @@ namespace URLNotesGrabberCORE
|
||||
return _cachedDbPath;
|
||||
}
|
||||
|
||||
#region IsActive
|
||||
|
||||
// Posts.IsActive and Notes.IsActive mean the same thing Blogs.IsActive does:
|
||||
// 0 = removed elsewhere (Rolodex), anything else (including NULL) = live.
|
||||
//
|
||||
// This crawler is a reader of all three. It never writes any IsActive column --
|
||||
// no INSERT lists it, no UPDATE sets it, and MapPrefixToColumn cannot map to it --
|
||||
// so a row removed in Rolodex is never resurrected by a re-crawl.
|
||||
//
|
||||
// Unlike Blogs.IsActive, the Posts and Notes columns are optional: they are added
|
||||
// from outside this app and are absent from databases that predate them. Naming a
|
||||
// missing column is a hard SQLite error ("no such column"), so every read asks the
|
||||
// schema first and simply drops the filter when the column is not there. The answer
|
||||
// is cached per database path, so adding the columns to a live database takes effect
|
||||
// on the next run.
|
||||
private static readonly Dictionary<string, bool> _isActiveColumnCache =
|
||||
new Dictionary<string, bool>(StringComparer.OrdinalIgnoreCase);
|
||||
private static readonly object _isActiveColumnLock = new object();
|
||||
|
||||
private static bool HasIsActiveColumn(string table, string? DBPath)
|
||||
{
|
||||
DBPath ??= GetDefaultDbPath();
|
||||
string cacheKey = DBPath + "|" + table;
|
||||
|
||||
lock (_isActiveColumnLock)
|
||||
{
|
||||
if (_isActiveColumnCache.TryGetValue(cacheKey, out bool cached))
|
||||
return cached;
|
||||
}
|
||||
|
||||
bool exists = false;
|
||||
try
|
||||
{
|
||||
using SQLiteConnection connection = new SQLiteConnection("Data Source=" + DBPath);
|
||||
connection.Open();
|
||||
|
||||
using SQLiteCommand command = new SQLiteCommand($"PRAGMA table_info({table});", connection);
|
||||
using SQLiteDataReader reader = command.ExecuteReader();
|
||||
while (reader.Read())
|
||||
{
|
||||
if (reader.GetString(1).Equals("IsActive", StringComparison.OrdinalIgnoreCase))
|
||||
{
|
||||
exists = true;
|
||||
break;
|
||||
}
|
||||
}
|
||||
}
|
||||
catch (Exception ex)
|
||||
{
|
||||
// Breakpoint here
|
||||
// An unreadable schema is treated as "no column" so the caller's query still runs.
|
||||
Console.WriteLine($"Error checking {table}.IsActive column: {ex.Message}");
|
||||
}
|
||||
|
||||
lock (_isActiveColumnLock)
|
||||
{
|
||||
_isActiveColumnCache[cacheKey] = exists;
|
||||
}
|
||||
|
||||
return exists;
|
||||
}
|
||||
|
||||
/// <summary>
|
||||
/// " AND COALESCE(alias.IsActive, 1) = 1" when the table carries the column, "" when it
|
||||
/// does not. NULL is read as live, the same way Rolodex reads Blogs.IsActive.
|
||||
/// </summary>
|
||||
private static string AndIsActive(string table, string alias = "", string? DBPath = null)
|
||||
{
|
||||
if (!HasIsActiveColumn(table, DBPath)) return string.Empty;
|
||||
|
||||
string qualifier = string.IsNullOrEmpty(alias) ? string.Empty : alias + ".";
|
||||
return $" AND COALESCE({qualifier}IsActive, 1) = 1";
|
||||
}
|
||||
|
||||
/// <summary>
|
||||
/// Same filter as <see cref="AndIsActive"/>, for a query that has no WHERE clause yet.
|
||||
/// </summary>
|
||||
private static string WhereIsActive(string table, string alias = "", string? DBPath = null)
|
||||
{
|
||||
string clause = AndIsActive(table, alias, DBPath);
|
||||
return clause.Length == 0 ? string.Empty : " WHERE" + clause.Substring(" AND".Length);
|
||||
}
|
||||
|
||||
#endregion IsActive
|
||||
|
||||
public static string Q(string input)
|
||||
{
|
||||
return "'" + input.Replace("'", "''") + "'";
|
||||
@@ -460,6 +553,8 @@ namespace URLNotesGrabberCORE
|
||||
{
|
||||
if (ownsConnection) connection.Open();
|
||||
|
||||
// IsActive is deliberately absent from this column list: a post removed
|
||||
// elsewhere must stay removed, so the crawler never writes that flag.
|
||||
string sql = @"INSERT INTO Posts (
|
||||
BlogName,
|
||||
PostID,
|
||||
@@ -507,7 +602,7 @@ namespace URLNotesGrabberCORE
|
||||
{
|
||||
if (ownsConnection) connection.Open();
|
||||
|
||||
string updateSql = "UPDATE Posts SET hasImage = @hasImage, DateModified = @DateModified WHERE blogName = @blogName AND postID = @postID";
|
||||
string updateSql = "UPDATE Posts SET hasImage = @hasImage, DateModified = @DateModified WHERE blogName = @blogName AND postID = @postID AND IFNULL(hasImage, 0) <> @hasImage";
|
||||
using SQLiteCommand updateCommand = new SQLiteCommand(updateSql, connection);
|
||||
updateCommand.Parameters.AddWithValue("@hasImage", hasImage ? 1 : 0);
|
||||
updateCommand.Parameters.AddWithValue("@DateModified", DateTime.Now.ToString("yyyy-MM-dd HH:mm:ss"));
|
||||
@@ -532,16 +627,17 @@ namespace URLNotesGrabberCORE
|
||||
}
|
||||
}
|
||||
|
||||
// Only update HasBeenOutput and DateAdded if a new post was inserted
|
||||
// Only reopen the blog for output if a new post was inserted. DateAdded records
|
||||
// when the blog first entered the registry and is never rewritten here -- a new
|
||||
// post is not a new blog.
|
||||
if (rowsInserted == 1)
|
||||
{
|
||||
try
|
||||
{
|
||||
string updateBlogSql = "UPDATE Blogs SET HasBeenOutput = 0, DateAdded = @DateAdded, DateModified = @DateModified WHERE BlogName = @BlogName";
|
||||
string updateBlogSql = "UPDATE Blogs SET HasBeenOutput = 0, DateModified = @DateModified WHERE BlogName = @BlogName AND (HasBeenOutput IS NULL OR HasBeenOutput <> 0)";
|
||||
using (var updateBlogCommand = new SQLiteCommand(updateBlogSql, connection))
|
||||
{
|
||||
updateBlogCommand.Parameters.AddWithValue("@BlogName", blogName);
|
||||
updateBlogCommand.Parameters.AddWithValue("@DateAdded", DateTime.Now.ToString("yyyy-MM-dd HH:mm:ss"));
|
||||
updateBlogCommand.Parameters.AddWithValue("@DateModified", DateTime.Now.ToString("yyyy-MM-dd HH:mm:ss"));
|
||||
updateBlogCommand.ExecuteNonQuery();
|
||||
}
|
||||
@@ -566,21 +662,40 @@ namespace URLNotesGrabberCORE
|
||||
|
||||
// Use INSERT OR IGNORE to avoid UNIQUE constraint errors when the date row already exists.
|
||||
// Also explicitly initialize APICount to 0 in case the table has no default.
|
||||
string sql = "INSERT OR IGNORE INTO DailyAPICount (Date, APICount, DateCreated) values(@date, 0, @DateCreated)";
|
||||
//
|
||||
// DailyAPICount is (Date TEXT PK, APICount INTEGER) — the crawler never creates or
|
||||
// migrates this table, and no code reads a creation timestamp off it, so the insert
|
||||
// names only those two columns. Naming a DateCreated column here used to throw
|
||||
// "no such column: DateCreated" into a silent catch, which meant the day's row was
|
||||
// never created and the tally sat at 0 for months.
|
||||
string sql = "INSERT OR IGNORE INTO DailyAPICount (Date, APICount) values(@date, 0)";
|
||||
using (SQLiteCommand command = new SQLiteCommand(sql, connection))
|
||||
{
|
||||
command.Parameters.AddWithValue("@date", DateTime.Today.ToShortDateString());
|
||||
command.Parameters.AddWithValue("@DateCreated", DateTime.Now.ToString("yyyy-MM-dd HH:mm:ss"));
|
||||
command.ExecuteNonQuery();
|
||||
}
|
||||
}
|
||||
catch (Exception ex)
|
||||
{
|
||||
// Breakpoint here
|
||||
//Console.WriteLine(ex.Message);
|
||||
ReportAPICountFailure($"Error creating the row for {DateTime.Today.ToShortDateString()}: {ex.Message}");
|
||||
}
|
||||
}
|
||||
|
||||
// Bookkeeping writes that fail identically on every API call would flood the console, but
|
||||
// swallowing them entirely is what hid the DateCreated bug. Log each distinct message once.
|
||||
// Messages embed today's date, so a date rollover reports afresh.
|
||||
private static void ReportAPICountFailure(string message)
|
||||
{
|
||||
lock (_apiCountFailureLock)
|
||||
{
|
||||
if (!_apiCountFailuresLogged.Add(message))
|
||||
return;
|
||||
}
|
||||
|
||||
Console.WriteLine($"[DailyAPICount] {message}");
|
||||
}
|
||||
|
||||
public static bool AddNote(string rootBlogName, string noteBlogName, long postID, long timestamp, string type, string? DBPath = null)
|
||||
{
|
||||
DBPath ??= GetDefaultDbPath();
|
||||
@@ -593,6 +708,8 @@ namespace URLNotesGrabberCORE
|
||||
{
|
||||
connection2.Open();
|
||||
|
||||
// INSERT OR IGNORE, and no IsActive in the column list: re-crawling a note
|
||||
// that was removed elsewhere leaves the existing row -- and its flag -- alone.
|
||||
string sql = "INSERT OR IGNORE INTO Notes (rootBlogName, noteBlogName, PostID, TimeStamp, Type, DatetimeCrawled, DateModified, DateCreated) values(@rootBlogName, @noteBlogName, @PostID, @TimeStamp, @Type, @DatetimeCrawled, @DateModified, @DateCreated)";
|
||||
using (SQLiteCommand command = new SQLiteCommand(sql, connection2))
|
||||
{
|
||||
@@ -625,7 +742,9 @@ namespace URLNotesGrabberCORE
|
||||
{
|
||||
try
|
||||
{
|
||||
string updateSql = "UPDATE Blogs SET HasBeenOutput = 0, DateModified = @DateModified WHERE BlogName = @BlogName";
|
||||
// HasBeenOutput IS NULL still counts as a change: the selection queries
|
||||
// test HasBeenOutput = 0, which a NULL would never match.
|
||||
string updateSql = "UPDATE Blogs SET HasBeenOutput = 0, DateModified = @DateModified WHERE BlogName = @BlogName AND (HasBeenOutput IS NULL OR HasBeenOutput <> 0)";
|
||||
using (var updateCommand = new SQLiteCommand(updateSql, connection2))
|
||||
{
|
||||
updateCommand.Parameters.AddWithValue("@BlogName", noteBlogName);
|
||||
@@ -690,7 +809,7 @@ namespace URLNotesGrabberCORE
|
||||
" P.HasNotesGathered," + Environment.NewLine +
|
||||
" P.NotFound," + Environment.NewLine +
|
||||
" P.PostDate" + Environment.NewLine +
|
||||
" FROM Posts P" + Environment.NewLine +
|
||||
" FROM Posts P" + WhereIsActive("Posts", "P", DBPath) + Environment.NewLine +
|
||||
")," + Environment.NewLine +
|
||||
"Unioned AS" + Environment.NewLine +
|
||||
"(" + Environment.NewLine +
|
||||
@@ -742,8 +861,8 @@ namespace URLNotesGrabberCORE
|
||||
" LEFT OUTER JOIN " + Environment.NewLine +
|
||||
" Notes ON Notes.RootBlogName = Posts.BlogName AND Notes.PostID = Posts.PostID " + Environment.NewLine +
|
||||
" LEFT OUTER JOIN " + Environment.NewLine +
|
||||
" ( select BlogName, count(PostID) as CNT from Posts group by BlogName) CNT on CNT.blogName = Posts.BlogName " +
|
||||
"WHERE NotFound = 0 " + Environment.NewLine;
|
||||
" ( select BlogName, count(PostID) as CNT from Posts" + WhereIsActive("Posts", "", DBPath) + " group by BlogName) CNT on CNT.blogName = Posts.BlogName " +
|
||||
"WHERE NotFound = 0 " + AndIsActive("Posts", "Posts", DBPath) + Environment.NewLine;
|
||||
|
||||
if (beforeDate.HasValue)
|
||||
{
|
||||
@@ -852,7 +971,7 @@ namespace URLNotesGrabberCORE
|
||||
{
|
||||
connection.Open();
|
||||
|
||||
string sql = "SELECT distinct RootBlogName as blogName, postID FROM Notes WHERE Notes.type = 'reply' order by RootBlogName, PostID";
|
||||
string sql = "SELECT distinct RootBlogName as blogName, postID FROM Notes WHERE Notes.type = 'reply'" + AndIsActive("Notes", "Notes", DBPath) + " order by RootBlogName, PostID";
|
||||
|
||||
using (SQLiteCommand command = new SQLiteCommand(sql, connection))
|
||||
{
|
||||
@@ -896,7 +1015,7 @@ namespace URLNotesGrabberCORE
|
||||
MAX(Notes.timestamp) as LatestTimestamp
|
||||
FROM Notes
|
||||
WHERE Notes.type = 'reply'
|
||||
AND (Notes.replyText IS NULL OR Notes.replyText = '' OR Notes.replyText = '.')
|
||||
AND (Notes.replyText IS NULL OR Notes.replyText = '' OR Notes.replyText = '.')" + AndIsActive("Notes", "Notes", DBPath) + @"
|
||||
GROUP BY Notes.RootBlogName, Notes.PostID
|
||||
ORDER BY LatestTimestamp ASC
|
||||
LIMIT @limit";
|
||||
@@ -944,7 +1063,7 @@ namespace URLNotesGrabberCORE
|
||||
INNER JOIN Notes N ON N.PostID = P.PostID AND N.RootBlogName = P.BlogName
|
||||
WHERE P.NotFound = 0
|
||||
AND N.type = 'reply'
|
||||
AND (N.replyText IS NULL OR N.replyText = '' OR N.replyText = '.')
|
||||
AND (N.replyText IS NULL OR N.replyText = '' OR N.replyText = '.')" + AndIsActive("Posts", "P", DBPath) + AndIsActive("Notes", "N", DBPath) + @"
|
||||
GROUP BY P.BlogName, P.PostID
|
||||
ORDER BY LatestTimestamp ASC";
|
||||
|
||||
@@ -993,7 +1112,9 @@ namespace URLNotesGrabberCORE
|
||||
using SQLiteConnection connection = new SQLiteConnection("Data Source=" + DBPath);
|
||||
int count = 0;
|
||||
|
||||
try { AddAPICount(); } catch { }
|
||||
// AddAPICount reports its own failures; this guard only stops a connection-level
|
||||
// problem from taking down the read below.
|
||||
try { AddAPICount(); } catch (Exception ex) { ReportAPICountFailure($"AddAPICount failed: {ex.Message}"); }
|
||||
|
||||
try
|
||||
{
|
||||
@@ -1001,6 +1122,7 @@ namespace URLNotesGrabberCORE
|
||||
|
||||
string sql = "SELECT APICount FROM DailyAPICount WHERE [Date] = @date";
|
||||
|
||||
bool rowFound = false;
|
||||
using (SQLiteCommand command = new SQLiteCommand(sql, connection))
|
||||
{
|
||||
command.Parameters.AddWithValue("@date", DateTime.Today.ToShortDateString());
|
||||
@@ -1008,10 +1130,16 @@ namespace URLNotesGrabberCORE
|
||||
{
|
||||
while (reader.Read())
|
||||
{
|
||||
rowFound = true;
|
||||
count = reader.GetInt32(0); // Assuming Id is the first column
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// A missing row means AddAPICount did not take. Returning a silent 0 here is what
|
||||
// made the tally look merely idle rather than broken.
|
||||
if (!rowFound)
|
||||
ReportAPICountFailure($"No row for {DateTime.Today.ToShortDateString()} after AddAPICount - reported count of 0 is not a real tally.");
|
||||
}
|
||||
catch (Exception ex)
|
||||
{
|
||||
@@ -1054,7 +1182,7 @@ namespace URLNotesGrabberCORE
|
||||
INNER JOIN Notes N ON N.NoteBlogName = B.BlogName
|
||||
WHERE N.TimeStamp >= 1535778000
|
||||
AND N.rootBlogName = B.BlogName
|
||||
AND B.IsActive = 1
|
||||
AND B.IsActive = 1" + AndIsActive("Notes", "N", DBPath) + @"
|
||||
GROUP BY B.BlogName
|
||||
ORDER BY MIN(N.Timestamp);";
|
||||
}
|
||||
@@ -1069,7 +1197,7 @@ namespace URLNotesGrabberCORE
|
||||
INNER JOIN Notes N ON N.NoteBlogName = B.BlogName
|
||||
WHERE N.TimeStamp >= 1535778000
|
||||
AND N.rootBlogName = B.BlogName
|
||||
AND B.IsActive = 1
|
||||
AND B.IsActive = 1" + AndIsActive("Notes", "N", DBPath) + @"
|
||||
AND (
|
||||
B.LikesPulled = 0
|
||||
OR COALESCE(B.LikesLastRefreshed, 0)
|
||||
@@ -1118,9 +1246,9 @@ namespace URLNotesGrabberCORE
|
||||
connection.Open();
|
||||
string sql = "";
|
||||
if (reblogsOnly)
|
||||
sql = "SELECT NoteBlogName as blogName, count(*) FROM notes INNER JOIN blogs ON blogs.BlogName = notes.NoteBlogName WHERE blogs.IsActive = @isActive AND type IN ('reblog', 'reply', 'posted') AND HasBeenOutput = 0 GROUP BY NoteBlogName ORDER BY count(*) DESC, BlogName LIMIT @top";
|
||||
sql = "SELECT NoteBlogName as blogName, count(*) FROM notes INNER JOIN blogs ON blogs.BlogName = notes.NoteBlogName WHERE blogs.IsActive = @isActive" + AndIsActive("Notes", "notes", DBPath) + " AND type IN ('reblog', 'reply', 'posted') AND HasBeenOutput = 0 GROUP BY NoteBlogName ORDER BY count(*) DESC, BlogName LIMIT @top";
|
||||
else
|
||||
sql = "SELECT NoteBlogName as blogName, count(*) FROM notes INNER JOIN blogs ON blogs.BlogName = notes.NoteBlogName WHERE blogs.IsActive = @isActive AND type NOT IN ('reblog', 'reply', 'posted') AND HasBeenOutput = 0 GROUP BY NoteBlogName ORDER BY count(*) DESC, BlogName LIMIT @top";
|
||||
sql = "SELECT NoteBlogName as blogName, count(*) FROM notes INNER JOIN blogs ON blogs.BlogName = notes.NoteBlogName WHERE blogs.IsActive = @isActive" + AndIsActive("Notes", "notes", DBPath) + " AND type NOT IN ('reblog', 'reply', 'posted') AND HasBeenOutput = 0 GROUP BY NoteBlogName ORDER BY count(*) DESC, BlogName LIMIT @top";
|
||||
|
||||
using (SQLiteCommand command = new SQLiteCommand(sql, connection))
|
||||
{
|
||||
@@ -1159,9 +1287,9 @@ namespace URLNotesGrabberCORE
|
||||
connection.Open();
|
||||
string sql = "";
|
||||
if (reblogsOnly)
|
||||
sql = "SELECT NoteBlogName as blogName, count(*) FROM notes INNER JOIN blogs ON blogs.BlogName = notes.NoteBlogName WHERE blogs.IsActive = @isActive AND type IN ('reblog', 'reply', 'posted') AND HasBeenOutput = 0 GROUP BY NoteBlogName ORDER BY count(*) DESC, BlogName LIMIT @top";
|
||||
sql = "SELECT NoteBlogName as blogName, count(*) FROM notes INNER JOIN blogs ON blogs.BlogName = notes.NoteBlogName WHERE blogs.IsActive = @isActive" + AndIsActive("Notes", "notes", DBPath) + " AND type IN ('reblog', 'reply', 'posted') AND HasBeenOutput = 0 GROUP BY NoteBlogName ORDER BY count(*) DESC, BlogName LIMIT @top";
|
||||
else
|
||||
sql = "SELECT NoteBlogName as blogName, count(*) FROM notes INNER JOIN blogs ON blogs.BlogName = notes.NoteBlogName WHERE blogs.IsActive = @isActive AND HasBeenOutput = 0 GROUP BY NoteBlogName ORDER BY count(*) DESC, BlogName LIMIT @top";
|
||||
sql = "SELECT NoteBlogName as blogName, count(*) FROM notes INNER JOIN blogs ON blogs.BlogName = notes.NoteBlogName WHERE blogs.IsActive = @isActive" + AndIsActive("Notes", "notes", DBPath) + " AND HasBeenOutput = 0 GROUP BY NoteBlogName ORDER BY count(*) DESC, BlogName LIMIT @top";
|
||||
|
||||
using (SQLiteCommand command = new SQLiteCommand(sql, connection))
|
||||
{
|
||||
@@ -1193,7 +1321,7 @@ namespace URLNotesGrabberCORE
|
||||
DBPath ??= GetDefaultDbPath();
|
||||
using SQLiteConnection connection = new SQLiteConnection("Data Source=" + DBPath);
|
||||
connection.Open();
|
||||
string sql = "SELECT BlogName, reblogURL, PostURL, Slug, ReblogKey, ReblogName, Summary, Quote, Body, Tags, Link, PhotoURL, PhotoCaption, DownloadedFiles, AudioCaption, Question, Answer, Title, RootBlogName, RootURL FROM Posts WHERE IFNULL(DownloadedFiles, '.') = '.'";
|
||||
string sql = "SELECT BlogName, reblogURL, PostURL, Slug, ReblogKey, ReblogName, Summary, Quote, Body, Tags, Link, PhotoURL, PhotoCaption, DownloadedFiles, AudioCaption, Question, Answer, Title, RootBlogName, RootURL FROM Posts WHERE IFNULL(DownloadedFiles, '.') = '.'" + AndIsActive("Posts", "", DBPath);
|
||||
using (SQLiteCommand command = new SQLiteCommand(sql, connection))
|
||||
{
|
||||
using (SQLiteDataReader reader = command.ExecuteReader())
|
||||
@@ -1235,7 +1363,10 @@ namespace URLNotesGrabberCORE
|
||||
connection.Open();
|
||||
|
||||
//string sql = "UPDATE Posts SET HasNotesGathered = 1, NotesGatheredDateTime = @notesGathered WHERE BlogName = @BlogName AND PostID = @PostID";
|
||||
string sql = "UPDATE Posts SET HasNotesGathered = 1, NotesGatheredDateTime = @notesGathered, DateModified = @dateModified WHERE PostID = @PostID AND (IFNULL(HasNotesGathered, 0) <> 1 OR IFNULL(NotesGatheredDateTime, 0) <> @notesGathered)";
|
||||
// NotesGatheredDateTime is crawl bookkeeping -- it moves on every pass and says
|
||||
// nothing about the post itself, so only the HasNotesGathered flag flipping
|
||||
// counts as a modification. The CASE reads the pre-UPDATE value of the flag.
|
||||
string sql = "UPDATE Posts SET HasNotesGathered = 1, NotesGatheredDateTime = @notesGathered, DateModified = CASE WHEN IFNULL(HasNotesGathered, 0) <> 1 THEN @dateModified ELSE DateModified END WHERE PostID = @PostID AND (IFNULL(HasNotesGathered, 0) <> 1 OR IFNULL(NotesGatheredDateTime, 0) <> @notesGathered)";
|
||||
using (SQLiteCommand command = new SQLiteCommand(sql, connection))
|
||||
{
|
||||
command.Parameters.AddWithValue("@notesGathered", DateTimeOffset.UtcNow.ToUnixTimeSeconds());
|
||||
@@ -1437,49 +1568,60 @@ namespace URLNotesGrabberCORE
|
||||
{
|
||||
if (ownsConnection) connection.Open();
|
||||
|
||||
// "." is TraverseDirectory/ReblogRecord's sentinel for "this field had no
|
||||
// matching line in this particular export file" -- not an empty value. A blog
|
||||
// with two export files in different formats (e.g. an "_2" duplicate folder, or
|
||||
// a Tumblr export whose field set changed over time) sends one record with a real
|
||||
// Title and another with Title = "." for the same PostID, and re-importing both
|
||||
// on every run must not let the "not supplied" record blank out what the other
|
||||
// one has. Every content field below is CASE-guarded the same way RootBlogName/
|
||||
// RootURL already were, and the change-detection ignores "." too so a "."-only
|
||||
// difference doesn't fire the UPDATE (and bump DateModified) on its own. Only "."
|
||||
// is treated as the sentinel -- an explicit empty string from a real field still
|
||||
// overwrites, same as before.
|
||||
string sql = "UPDATE Posts SET ";
|
||||
sql += "postDate = @postDate, ";
|
||||
sql += "reblogURL = @reblogURL, ";
|
||||
sql += "postURL = @postURL, ";
|
||||
sql += "slug = @slug, ";
|
||||
sql += "reblogKey = @reblogKey, ";
|
||||
sql += "reblogName = @reblogName, ";
|
||||
sql += "summary = @summary, ";
|
||||
sql += "quote = @quote, ";
|
||||
sql += "body = @body, ";
|
||||
sql += "tags = @tags, ";
|
||||
sql += "link = @link, ";
|
||||
sql += "photoURL = @photoURL, ";
|
||||
sql += "photoCaption = @photoCaption, ";
|
||||
sql += "downloadedFiles = @downloadedFiles, ";
|
||||
sql += "audioCaption = @audioCaption, ";
|
||||
sql += "question = @question, ";
|
||||
sql += "answer = @answer, ";
|
||||
sql += "title = @title, ";
|
||||
sql += "postDate = CASE WHEN @postDate = '.' THEN postDate ELSE @postDate END, ";
|
||||
sql += "reblogURL = CASE WHEN @reblogURL = '.' THEN reblogURL ELSE @reblogURL END, ";
|
||||
sql += "postURL = CASE WHEN @postURL = '.' THEN postURL ELSE @postURL END, ";
|
||||
sql += "slug = CASE WHEN @slug = '.' THEN slug ELSE @slug END, ";
|
||||
sql += "reblogKey = CASE WHEN @reblogKey = '.' THEN reblogKey ELSE @reblogKey END, ";
|
||||
sql += "reblogName = CASE WHEN @reblogName = '.' THEN reblogName ELSE @reblogName END, ";
|
||||
sql += "summary = CASE WHEN @summary = '.' THEN summary ELSE @summary END, ";
|
||||
sql += "quote = CASE WHEN @quote = '.' THEN quote ELSE @quote END, ";
|
||||
sql += "body = CASE WHEN @body = '.' THEN body ELSE @body END, ";
|
||||
sql += "tags = CASE WHEN @tags = '.' THEN tags ELSE @tags END, ";
|
||||
sql += "link = CASE WHEN @link = '.' THEN link ELSE @link END, ";
|
||||
sql += "photoURL = CASE WHEN @photoURL = '.' THEN photoURL ELSE @photoURL END, ";
|
||||
sql += "photoCaption = CASE WHEN @photoCaption = '.' THEN photoCaption ELSE @photoCaption END, ";
|
||||
sql += "downloadedFiles = CASE WHEN @downloadedFiles = '.' THEN downloadedFiles ELSE @downloadedFiles END, ";
|
||||
sql += "audioCaption = CASE WHEN @audioCaption = '.' THEN audioCaption ELSE @audioCaption END, ";
|
||||
sql += "question = CASE WHEN @question = '.' THEN question ELSE @question END, ";
|
||||
sql += "answer = CASE WHEN @answer = '.' THEN answer ELSE @answer END, ";
|
||||
sql += "title = CASE WHEN @title = '.' THEN title ELSE @title END, ";
|
||||
sql += "DateModified = @dateModified, ";
|
||||
sql += "RootBlogName = CASE WHEN @rootBlogName IS NULL OR @rootBlogName = '' OR @rootBlogName = '.' THEN RootBlogName ELSE @rootBlogName END, ";
|
||||
sql += "RootURL = CASE WHEN @rootURL IS NULL OR @rootURL = '' OR @rootURL = '.' THEN RootURL ELSE @rootURL END, ";
|
||||
sql += "hasImage = @hasImage, ";
|
||||
sql += "ByLikes = MAX(IFNULL(ByLikes, 0), @byLikes) ";
|
||||
sql += " WHERE BlogName = @BlogName AND PostID = @PostID AND (";
|
||||
sql += "IFNULL(postDate, '') <> @postDate OR ";
|
||||
sql += "IFNULL(reblogURL, '') <> @reblogURL OR ";
|
||||
sql += "IFNULL(postURL, '') <> @postURL OR ";
|
||||
sql += "IFNULL(slug, '') <> @slug OR ";
|
||||
sql += "IFNULL(reblogKey, '') <> @reblogKey OR ";
|
||||
sql += "IFNULL(reblogName, '') <> @reblogName OR ";
|
||||
sql += "IFNULL(summary, '') <> @summary OR ";
|
||||
sql += "IFNULL(quote, '') <> @quote OR ";
|
||||
sql += "IFNULL(body, '') <> @body OR ";
|
||||
sql += "IFNULL(tags, '') <> @tags OR ";
|
||||
sql += "IFNULL(link, '') <> @link OR ";
|
||||
sql += "IFNULL(photoURL, '') <> @photoURL OR ";
|
||||
sql += "IFNULL(photoCaption, '') <> @photoCaption OR ";
|
||||
sql += "IFNULL(downloadedFiles, '') <> @downloadedFiles OR ";
|
||||
sql += "IFNULL(audioCaption, '') <> @audioCaption OR ";
|
||||
sql += "IFNULL(question, '') <> @question OR ";
|
||||
sql += "IFNULL(answer, '') <> @answer OR ";
|
||||
sql += "IFNULL(title, '') <> @title OR ";
|
||||
sql += "(@postDate <> '.' AND IFNULL(postDate, '') <> @postDate) OR ";
|
||||
sql += "(@reblogURL <> '.' AND IFNULL(reblogURL, '') <> @reblogURL) OR ";
|
||||
sql += "(@postURL <> '.' AND IFNULL(postURL, '') <> @postURL) OR ";
|
||||
sql += "(@slug <> '.' AND IFNULL(slug, '') <> @slug) OR ";
|
||||
sql += "(@reblogKey <> '.' AND IFNULL(reblogKey, '') <> @reblogKey) OR ";
|
||||
sql += "(@reblogName <> '.' AND IFNULL(reblogName, '') <> @reblogName) OR ";
|
||||
sql += "(@summary <> '.' AND IFNULL(summary, '') <> @summary) OR ";
|
||||
sql += "(@quote <> '.' AND IFNULL(quote, '') <> @quote) OR ";
|
||||
sql += "(@body <> '.' AND IFNULL(body, '') <> @body) OR ";
|
||||
sql += "(@tags <> '.' AND IFNULL(tags, '') <> @tags) OR ";
|
||||
sql += "(@link <> '.' AND IFNULL(link, '') <> @link) OR ";
|
||||
sql += "(@photoURL <> '.' AND IFNULL(photoURL, '') <> @photoURL) OR ";
|
||||
sql += "(@photoCaption <> '.' AND IFNULL(photoCaption, '') <> @photoCaption) OR ";
|
||||
sql += "(@downloadedFiles <> '.' AND IFNULL(downloadedFiles, '') <> @downloadedFiles) OR ";
|
||||
sql += "(@audioCaption <> '.' AND IFNULL(audioCaption, '') <> @audioCaption) OR ";
|
||||
sql += "(@question <> '.' AND IFNULL(question, '') <> @question) OR ";
|
||||
sql += "(@answer <> '.' AND IFNULL(answer, '') <> @answer) OR ";
|
||||
sql += "(@title <> '.' AND IFNULL(title, '') <> @title) OR ";
|
||||
sql += "IFNULL(hasImage, 0) <> @hasImage OR ";
|
||||
sql += "(@byLikes = 1 AND IFNULL(ByLikes, 0) = 0) OR ";
|
||||
sql += "((@rootBlogName IS NOT NULL AND @rootBlogName <> '' AND @rootBlogName <> '.') AND IFNULL(RootBlogName, '') <> @rootBlogName) OR ";
|
||||
@@ -1592,7 +1734,8 @@ namespace URLNotesGrabberCORE
|
||||
string sql = @"UPDATE Blogs
|
||||
SET LikesNewestTimestamp = MAX(COALESCE(LikesNewestTimestamp, 0), @newest),
|
||||
DateModified = @modified
|
||||
WHERE BlogName = @name";
|
||||
WHERE BlogName = @name
|
||||
AND COALESCE(LikesNewestTimestamp, 0) < @newest";
|
||||
using (SQLiteCommand command = new SQLiteCommand(sql, connection))
|
||||
{
|
||||
command.Parameters.AddWithValue("@newest", newestTimestamp);
|
||||
@@ -1654,7 +1797,11 @@ namespace URLNotesGrabberCORE
|
||||
{
|
||||
command.Parameters.AddWithValue("@APICount", APICount);
|
||||
command.Parameters.AddWithValue("@date", DateTime.Today.ToShortDateString());
|
||||
command.ExecuteNonQuery();
|
||||
|
||||
// No row for today means this UPDATE matched nothing and the increment was
|
||||
// thrown away, while the value returned below still looks like a real count.
|
||||
if (command.ExecuteNonQuery() == 0)
|
||||
ReportAPICountFailure($"UPDATE matched no row for {DateTime.Today.ToShortDateString()} - the count of {APICount} was not persisted.");
|
||||
}
|
||||
}
|
||||
catch (Exception ex)
|
||||
@@ -1684,7 +1831,7 @@ namespace URLNotesGrabberCORE
|
||||
//string sql = "UPDATE Notes SET replyText = @replyText WHERE rootBlogName = @rootBlogName AND PostID = @PostID AND noteBlogName = @noteBlogName AND TimeStamp = @TimeStamp AND Type = 'reply'";
|
||||
// Match on (noteBlogName, TimeStamp ±5s) only - a reply by a given blog at a given timestamp is the same reply across the original post and every reblog of it, so this fans out across reblog chains in one shot. Tolerance absorbs the ~1s drift between what -collect stored and what mode=conversation returns now.
|
||||
// Only fan out to rows that match the SELECT criteria in GetRepliesWithFilledText (NULL/empty/legacy-'.'). Never overwrite '?' (confirmed-empty) or already-fetched text.
|
||||
string sql = "UPDATE Notes SET replyText = @replyText, DateModified = @dateModified WHERE noteBlogName = @noteBlogName AND ABS(TimeStamp - @TimeStamp) <= 5 AND Type = 'reply' AND (replyText IS NULL OR replyText = '' OR replyText = '.')";
|
||||
string sql = "UPDATE Notes SET replyText = @replyText, DateModified = @dateModified WHERE noteBlogName = @noteBlogName AND ABS(TimeStamp - @TimeStamp) <= 5 AND Type = 'reply' AND (replyText IS NULL OR replyText = '' OR replyText = '.') AND (replyText IS NULL OR replyText <> @replyText)";
|
||||
using (SQLiteCommand command = new SQLiteCommand(sql, connection))
|
||||
{
|
||||
command.Parameters.AddWithValue("@replyText", replyText ?? "?");
|
||||
@@ -1859,6 +2006,8 @@ namespace URLNotesGrabberCORE
|
||||
{
|
||||
if (ownsConnection) connection.Open();
|
||||
|
||||
// As in AddPost, IsActive is never written -- neither here nor in the
|
||||
// UPDATE below, which is why an ingest cannot un-remove a post.
|
||||
string insertSql = @"INSERT INTO Posts (
|
||||
BlogName, PostID, reblogURL, PostDate, PostURL, Slug,
|
||||
ReblogKey, ReblogName, Summary, Quote, Body, Tags, Link,
|
||||
@@ -1917,29 +2066,65 @@ namespace URLNotesGrabberCORE
|
||||
|
||||
if (rowsInserted == 0)
|
||||
{
|
||||
// NULL is this function's sentinel for "this file's record had no line for
|
||||
// that field" (IngestMode's G(key) misses return null; LegacyPostsDbImporter
|
||||
// passes null straight from a NULL source column) -- it does not mean "clear
|
||||
// this field". --ingest's entire reason to exist is reconciling multiple
|
||||
// export files for the same (BlogName, PostID) -- IngestMode normalizes a
|
||||
// "_2"-suffixed duplicate folder onto the same blog name specifically so a
|
||||
// second, differently-formatted file for a post it already has gets merged in.
|
||||
// Files are walked in filesystem enumeration order, not sorted, so which
|
||||
// file's UpsertPostFromTextFile call runs last for a given PostID is
|
||||
// effectively arbitrary. An unconditional SET here would let whichever file
|
||||
// processed last silently null out every column its own record didn't carry,
|
||||
// erasing real content the other file had -- the opposite of "clean up". Each
|
||||
// column is CASE-guarded to keep the existing value when this call's parameter
|
||||
// is NULL, and the change-detection ignores a NULL-vs-real mismatch the same
|
||||
// way, so a partial record converges into the row instead of overwriting it.
|
||||
string updateSql = @"UPDATE Posts SET
|
||||
reblogURL = @reblogURL,
|
||||
PostDate = @PostDate,
|
||||
PostURL = @PostURL,
|
||||
Slug = @Slug,
|
||||
ReblogKey = @ReblogKey,
|
||||
ReblogName = @ReblogName,
|
||||
Summary = @Summary,
|
||||
Quote = @Quote,
|
||||
Body = @Body,
|
||||
Tags = @Tags,
|
||||
Link = @Link,
|
||||
PhotoURL = @PhotoURL,
|
||||
PhotoCaption = @PhotoCaption,
|
||||
DownloadedFiles = @DownloadedFiles,
|
||||
AudioCaption = @AudioCaption,
|
||||
Question = @Question,
|
||||
Answer = @Answer,
|
||||
Title = @Title,
|
||||
PostType = @PostType,
|
||||
reblogURL = CASE WHEN @reblogURL IS NULL THEN reblogURL ELSE @reblogURL END,
|
||||
PostDate = CASE WHEN @PostDate IS NULL THEN PostDate ELSE @PostDate END,
|
||||
PostURL = CASE WHEN @PostURL IS NULL THEN PostURL ELSE @PostURL END,
|
||||
Slug = CASE WHEN @Slug IS NULL THEN Slug ELSE @Slug END,
|
||||
ReblogKey = CASE WHEN @ReblogKey IS NULL THEN ReblogKey ELSE @ReblogKey END,
|
||||
ReblogName = CASE WHEN @ReblogName IS NULL THEN ReblogName ELSE @ReblogName END,
|
||||
Summary = CASE WHEN @Summary IS NULL THEN Summary ELSE @Summary END,
|
||||
Quote = CASE WHEN @Quote IS NULL THEN Quote ELSE @Quote END,
|
||||
Body = CASE WHEN @Body IS NULL THEN Body ELSE @Body END,
|
||||
Tags = CASE WHEN @Tags IS NULL THEN Tags ELSE @Tags END,
|
||||
Link = CASE WHEN @Link IS NULL THEN Link ELSE @Link END,
|
||||
PhotoURL = CASE WHEN @PhotoURL IS NULL THEN PhotoURL ELSE @PhotoURL END,
|
||||
PhotoCaption = CASE WHEN @PhotoCaption IS NULL THEN PhotoCaption ELSE @PhotoCaption END,
|
||||
DownloadedFiles = CASE WHEN @DownloadedFiles IS NULL THEN DownloadedFiles ELSE @DownloadedFiles END,
|
||||
AudioCaption = CASE WHEN @AudioCaption IS NULL THEN AudioCaption ELSE @AudioCaption END,
|
||||
Question = CASE WHEN @Question IS NULL THEN Question ELSE @Question END,
|
||||
Answer = CASE WHEN @Answer IS NULL THEN Answer ELSE @Answer END,
|
||||
Title = CASE WHEN @Title IS NULL THEN Title ELSE @Title END,
|
||||
PostType = CASE WHEN @PostType IS NULL THEN PostType ELSE @PostType END,
|
||||
HasImage = @HasImage,
|
||||
DateModified = @DateModified
|
||||
WHERE BlogName = @BlogName AND PostID = @PostID";
|
||||
WHERE BlogName = @BlogName AND PostID = @PostID AND (
|
||||
(@reblogURL IS NOT NULL AND IFNULL(reblogURL, '') <> @reblogURL) OR
|
||||
(@PostDate IS NOT NULL AND IFNULL(PostDate, '') <> @PostDate) OR
|
||||
(@PostURL IS NOT NULL AND IFNULL(PostURL, '') <> @PostURL) OR
|
||||
(@Slug IS NOT NULL AND IFNULL(Slug, '') <> @Slug) OR
|
||||
(@ReblogKey IS NOT NULL AND IFNULL(ReblogKey, '') <> @ReblogKey) OR
|
||||
(@ReblogName IS NOT NULL AND IFNULL(ReblogName, '') <> @ReblogName) OR
|
||||
(@Summary IS NOT NULL AND IFNULL(Summary, '') <> @Summary) OR
|
||||
(@Quote IS NOT NULL AND IFNULL(Quote, '') <> @Quote) OR
|
||||
(@Body IS NOT NULL AND IFNULL(Body, '') <> @Body) OR
|
||||
(@Tags IS NOT NULL AND IFNULL(Tags, '') <> @Tags) OR
|
||||
(@Link IS NOT NULL AND IFNULL(Link, '') <> @Link) OR
|
||||
(@PhotoURL IS NOT NULL AND IFNULL(PhotoURL, '') <> @PhotoURL) OR
|
||||
(@PhotoCaption IS NOT NULL AND IFNULL(PhotoCaption, '') <> @PhotoCaption) OR
|
||||
(@DownloadedFiles IS NOT NULL AND IFNULL(DownloadedFiles, '') <> @DownloadedFiles) OR
|
||||
(@AudioCaption IS NOT NULL AND IFNULL(AudioCaption, '') <> @AudioCaption) OR
|
||||
(@Question IS NOT NULL AND IFNULL(Question, '') <> @Question) OR
|
||||
(@Answer IS NOT NULL AND IFNULL(Answer, '') <> @Answer) OR
|
||||
(@Title IS NOT NULL AND IFNULL(Title, '') <> @Title) OR
|
||||
(@PostType IS NOT NULL AND IFNULL(PostType, '') <> @PostType) OR
|
||||
IFNULL(HasImage, 0) <> @HasImage
|
||||
)";
|
||||
|
||||
using (var cmd = new SQLiteCommand(updateSql, connection))
|
||||
{
|
||||
@@ -1989,7 +2174,7 @@ namespace URLNotesGrabberCORE
|
||||
PhotoURL, PhotoCaption, DownloadedFiles, AudioCaption,
|
||||
Question, Answer, Title, PostType,
|
||||
HasImage, DateCreated, DateModified
|
||||
FROM Posts WHERE BlogName = @BlogName";
|
||||
FROM Posts WHERE BlogName = @BlogName" + AndIsActive("Posts", "", DBPath);
|
||||
using var cmd = new SQLiteCommand(sql, connection);
|
||||
cmd.Parameters.AddWithValue("@BlogName", blogName);
|
||||
using var reader = cmd.ExecuteReader();
|
||||
@@ -2038,7 +2223,7 @@ namespace URLNotesGrabberCORE
|
||||
PhotoURL, PhotoCaption, DownloadedFiles, AudioCaption,
|
||||
Question, Answer, Title, PostType,
|
||||
HasImage, DateCreated, DateModified
|
||||
FROM Posts WHERE BlogName = @BlogName AND PostID = @PostID";
|
||||
FROM Posts WHERE BlogName = @BlogName AND PostID = @PostID" + AndIsActive("Posts", "", DBPath);
|
||||
using var cmd = new SQLiteCommand(sql, connection);
|
||||
cmd.Parameters.AddWithValue("@BlogName", blogName);
|
||||
cmd.Parameters.AddWithValue("@PostID", postId);
|
||||
@@ -2088,7 +2273,7 @@ namespace URLNotesGrabberCORE
|
||||
PhotoURL, PhotoCaption, DownloadedFiles, AudioCaption,
|
||||
Question, Answer, Title, PostType,
|
||||
HasImage, DateCreated, DateModified
|
||||
FROM Posts WHERE PostID = @PostID LIMIT 1";
|
||||
FROM Posts WHERE PostID = @PostID" + AndIsActive("Posts", "", DBPath) + @" LIMIT 1";
|
||||
using var cmd = new SQLiteCommand(sql, connection);
|
||||
cmd.Parameters.AddWithValue("@PostID", postId);
|
||||
using var reader = cmd.ExecuteReader();
|
||||
@@ -2123,7 +2308,11 @@ namespace URLNotesGrabberCORE
|
||||
};
|
||||
}
|
||||
|
||||
public static void SetBlogTTFolderPath(string blogName, string? path, string? DBPath = null)
|
||||
// Returns true only when a row's TTFolderPath actually changed. A false means either
|
||||
// the row already held this value or no row matched the name -- callers must not
|
||||
// report a write they did not get, which is how a --updatepaths run could once print
|
||||
// "Updated <blog>" for every metadata file while leaving the column entirely NULL.
|
||||
public static bool SetBlogTTFolderPath(string blogName, string? path, string? DBPath = null)
|
||||
{
|
||||
DBPath ??= GetDefaultDbPath();
|
||||
try { AddBlog(blogName, false, DBPath); } catch { }
|
||||
@@ -2131,24 +2320,40 @@ namespace URLNotesGrabberCORE
|
||||
using var connection = new SQLiteConnection("Data Source=" + DBPath);
|
||||
connection.Open();
|
||||
using var cmd = new SQLiteCommand(
|
||||
"UPDATE Blogs SET TTFolderPath = @path, DateModified = @modified WHERE BlogName = @name",
|
||||
"UPDATE Blogs SET TTFolderPath = @path, DateModified = @modified WHERE BlogName = @name AND IFNULL(TTFolderPath, '') <> IFNULL(@path, '')",
|
||||
connection);
|
||||
cmd.Parameters.AddWithValue("@path", (object?)path ?? DBNull.Value);
|
||||
cmd.Parameters.AddWithValue("@modified", DateTime.Now.ToString("yyyy-MM-dd HH:mm:ss"));
|
||||
cmd.Parameters.AddWithValue("@name", blogName);
|
||||
cmd.ExecuteNonQuery();
|
||||
return cmd.ExecuteNonQuery() > 0;
|
||||
}
|
||||
|
||||
// Whether a Blogs row exists under this exact name. BlogName is a BINARY-collated
|
||||
// primary key, so a metadata filename that differs only in case is a different blog
|
||||
// as far as the UPDATE above is concerned -- worth telling the user about.
|
||||
public static bool BlogExists(string blogName, string? DBPath = null)
|
||||
{
|
||||
DBPath ??= GetDefaultDbPath();
|
||||
|
||||
using var connection = new SQLiteConnection("Data Source=" + DBPath);
|
||||
connection.Open();
|
||||
using var cmd = new SQLiteCommand("SELECT 1 FROM Blogs WHERE BlogName = @name", connection);
|
||||
cmd.Parameters.AddWithValue("@name", blogName);
|
||||
return cmd.ExecuteScalar() != null;
|
||||
}
|
||||
|
||||
// Partial UPDATE used by the correct-apply path. fieldsToUpdate maps
|
||||
// ThreeTxtFileHelper prefix names ("Reblog URL", "Body", etc.) to non-empty
|
||||
// values pulled from a BAK file. Only those columns + DateModified are written;
|
||||
// other content columns and all engagement columns are left intact.
|
||||
// Returns true if a row was matched (and therefore updated).
|
||||
// Returns true if a row was actually changed. A row whose columns already hold
|
||||
// the incoming values is left alone, DateModified included.
|
||||
public static bool UpdatePostContentFields(string blogName, string postId, IDictionary<string, string> fieldsToUpdate, string? DBPath = null)
|
||||
{
|
||||
DBPath ??= GetDefaultDbPath();
|
||||
|
||||
var setClauses = new List<string>();
|
||||
var changedClauses = new List<string>();
|
||||
var parameters = new List<(string Name, object Value)>();
|
||||
|
||||
foreach (var kvp in fieldsToUpdate)
|
||||
@@ -2158,6 +2363,7 @@ namespace URLNotesGrabberCORE
|
||||
if (column == null) continue;
|
||||
string paramName = "@p" + parameters.Count;
|
||||
setClauses.Add($"{column} = {paramName}");
|
||||
changedClauses.Add($"IFNULL({column}, '') <> {paramName}");
|
||||
parameters.Add((paramName, kvp.Value));
|
||||
}
|
||||
|
||||
@@ -2169,7 +2375,7 @@ namespace URLNotesGrabberCORE
|
||||
using var connection = new SQLiteConnection("Data Source=" + DBPath);
|
||||
connection.Open();
|
||||
|
||||
string sql = $"UPDATE Posts SET {string.Join(", ", setClauses)} WHERE BlogName = @BlogName AND PostID = @PostID";
|
||||
string sql = $"UPDATE Posts SET {string.Join(", ", setClauses)} WHERE BlogName = @BlogName AND PostID = @PostID AND ({string.Join(" OR ", changedClauses)})";
|
||||
using var cmd = new SQLiteCommand(sql, connection);
|
||||
foreach (var (name, value) in parameters)
|
||||
cmd.Parameters.AddWithValue(name, value);
|
||||
@@ -2207,24 +2413,48 @@ namespace URLNotesGrabberCORE
|
||||
};
|
||||
}
|
||||
|
||||
public static List<(string BlogName, string? TTFolderPath)> GetAllBlogsWithTTFolderPath(string? DBPath = null)
|
||||
// Export targets only: active blogs that actually carry a TTFolderPath.
|
||||
// Blogs is a 144k-row crawl registry and only the few hundred blogs downloaded
|
||||
// locally have a folder, so returning the unset rows made --output print a skip
|
||||
// line for every blog Tumblr has ever handed us.
|
||||
public static List<(string BlogName, string TTFolderPath)> GetAllBlogsWithTTFolderPath(string? DBPath = null)
|
||||
{
|
||||
DBPath ??= GetDefaultDbPath();
|
||||
var results = new List<(string, string?)>();
|
||||
var results = new List<(string, string)>();
|
||||
|
||||
using var connection = new SQLiteConnection("Data Source=" + DBPath);
|
||||
connection.Open();
|
||||
using var cmd = new SQLiteCommand("SELECT BlogName, TTFolderPath FROM Blogs WHERE IsActive = 1", connection);
|
||||
using var cmd = new SQLiteCommand(
|
||||
"SELECT BlogName, TRIM(TTFolderPath) FROM Blogs WHERE IsActive = 1 AND IFNULL(TRIM(TTFolderPath), '') <> '' ORDER BY BlogName",
|
||||
connection);
|
||||
using var reader = cmd.ExecuteReader();
|
||||
while (reader.Read())
|
||||
{
|
||||
string name = reader.GetString(0);
|
||||
string? path = reader.IsDBNull(1) ? null : reader.GetString(1);
|
||||
results.Add((name, path));
|
||||
}
|
||||
results.Add((reader.GetString(0), reader.GetString(1)));
|
||||
return results;
|
||||
}
|
||||
|
||||
// Companion counts for the messages --output and --updatepaths print about coverage.
|
||||
public static int CountActiveBlogs(string? DBPath = null)
|
||||
{
|
||||
DBPath ??= GetDefaultDbPath();
|
||||
|
||||
using var connection = new SQLiteConnection("Data Source=" + DBPath);
|
||||
connection.Open();
|
||||
using var cmd = new SQLiteCommand("SELECT COUNT(*) FROM Blogs WHERE IsActive = 1", connection);
|
||||
return Convert.ToInt32(cmd.ExecuteScalar());
|
||||
}
|
||||
|
||||
public static int CountBlogsWithTTFolderPath(string? DBPath = null)
|
||||
{
|
||||
DBPath ??= GetDefaultDbPath();
|
||||
|
||||
using var connection = new SQLiteConnection("Data Source=" + DBPath);
|
||||
connection.Open();
|
||||
using var cmd = new SQLiteCommand(
|
||||
"SELECT COUNT(*) FROM Blogs WHERE IFNULL(TRIM(TTFolderPath), '') <> ''", connection);
|
||||
return Convert.ToInt32(cmd.ExecuteScalar());
|
||||
}
|
||||
|
||||
private static string SafeStr(SQLiteDataReader reader, int ordinal)
|
||||
{
|
||||
return reader.IsDBNull(ordinal) ? string.Empty : reader.GetValue(ordinal)?.ToString() ?? string.Empty;
|
||||
|
||||
@@ -29,6 +29,8 @@ namespace URLNotesGrabberCORE
|
||||
Console.WriteLine($"Reading legacy posts.db: {legacyDbPath}");
|
||||
|
||||
int blogsCopied = 0;
|
||||
int blogPathsWritten = 0;
|
||||
int blogsWithoutPath = 0;
|
||||
int postsUpserted = 0;
|
||||
int errors = 0;
|
||||
|
||||
@@ -48,7 +50,13 @@ namespace URLNotesGrabberCORE
|
||||
if (string.IsNullOrWhiteSpace(blogName)) continue;
|
||||
try
|
||||
{
|
||||
DataAccess.SetBlogTTFolderPath(blogName, ttFolderPath);
|
||||
// A legacy row whose TTFolderPath was already NULL copies nothing.
|
||||
// Counting it as "copied" is what hid the fact that this import has
|
||||
// never populated a single path.
|
||||
if (string.IsNullOrWhiteSpace(ttFolderPath))
|
||||
blogsWithoutPath++;
|
||||
else if (DataAccess.SetBlogTTFolderPath(blogName, ttFolderPath.Trim()))
|
||||
blogPathsWritten++;
|
||||
blogsCopied++;
|
||||
}
|
||||
catch (Exception ex)
|
||||
@@ -58,7 +66,7 @@ namespace URLNotesGrabberCORE
|
||||
}
|
||||
}
|
||||
}
|
||||
Console.WriteLine($" Blogs copied: {blogsCopied}");
|
||||
Console.WriteLine($" Blogs seen: {blogsCopied}, TTFolderPath written: {blogPathsWritten}, legacy rows with no path: {blogsWithoutPath}");
|
||||
|
||||
// 2) Copy Posts
|
||||
try
|
||||
@@ -138,7 +146,8 @@ namespace URLNotesGrabberCORE
|
||||
}
|
||||
|
||||
Console.WriteLine($"\n========== Legacy import summary ==========");
|
||||
Console.WriteLine($"Blogs copied: {blogsCopied}");
|
||||
Console.WriteLine($"Blogs seen: {blogsCopied}");
|
||||
Console.WriteLine($"Paths written: {blogPathsWritten} (legacy rows with no path: {blogsWithoutPath})");
|
||||
Console.WriteLine($"Posts upserted: {postsUpserted}");
|
||||
Console.WriteLine($"Errors: {errors}");
|
||||
return errors == 0 ? 0 : 2;
|
||||
|
||||
@@ -8,28 +8,51 @@ namespace URLNotesGrabberCORE
|
||||
// field order). Reads from TL.db via DataAccess.GetAllPostsForBlog.
|
||||
public static class OutputMode
|
||||
{
|
||||
public static int Run(IConfiguration config)
|
||||
public static int Run(IConfiguration config, string[]? args = null)
|
||||
{
|
||||
DataAccess.EnsureTTFileHelperColumnsExist();
|
||||
|
||||
var blogs = DataAccess.GetAllBlogsWithTTFolderPath();
|
||||
Console.WriteLine($"Found {blogs.Count} blog(s) to process.");
|
||||
string dbPath = DataAccess.GetActiveDbPath();
|
||||
Console.WriteLine($"Database: {Path.GetFullPath(dbPath)}");
|
||||
|
||||
foreach (var (blogName, ttFolderPath) in blogs)
|
||||
if (!RefreshPaths(config, args ?? Array.Empty<string>()))
|
||||
return 1;
|
||||
|
||||
var blogs = DataAccess.GetAllBlogsWithTTFolderPath();
|
||||
int activeBlogs = DataAccess.CountActiveBlogs();
|
||||
Console.WriteLine($"{blogs.Count} of {activeBlogs} active blog(s) have a TTFolderPath.");
|
||||
|
||||
if (blogs.Count == 0)
|
||||
{
|
||||
Console.WriteLine($"\nNothing to export: no blog in {Path.GetFullPath(dbPath)} has a TTFolderPath.");
|
||||
Console.WriteLine("Point --output at a TumblThree root so it can populate them: --output <root>,");
|
||||
Console.WriteLine("or set appSettings:PathTTRoot so the refresh runs automatically.");
|
||||
return 1;
|
||||
}
|
||||
|
||||
int missingFolderCount = 0;
|
||||
int writtenCount = 0;
|
||||
|
||||
foreach (var (blogName, folder) in blogs)
|
||||
{
|
||||
Console.WriteLine($"\nProcessing blog: {blogName}");
|
||||
|
||||
if (string.IsNullOrWhiteSpace(ttFolderPath) || !Directory.Exists(ttFolderPath))
|
||||
// A stored path that this machine cannot see means the value was written on
|
||||
// another machine -- re-running --updatepaths locally is the fix, so say so
|
||||
// rather than lumping it in with "not set".
|
||||
if (!Directory.Exists(folder))
|
||||
{
|
||||
Console.WriteLine($" TTFolderPath does not exist or is not set. Skipping.");
|
||||
Console.WriteLine($" TTFolderPath folder not found: {folder}. Skipping.");
|
||||
missingFolderCount++;
|
||||
continue;
|
||||
}
|
||||
|
||||
Console.WriteLine($" TTFolderPath: {ttFolderPath}");
|
||||
Console.WriteLine($" TTFolderPath: {folder}");
|
||||
writtenCount++;
|
||||
|
||||
try
|
||||
{
|
||||
foreach (var bakFile in Directory.GetFiles(ttFolderPath, "*.bak"))
|
||||
foreach (var bakFile in Directory.GetFiles(folder, "*.bak"))
|
||||
File.Delete(bakFile);
|
||||
}
|
||||
catch (Exception ex)
|
||||
@@ -37,7 +60,7 @@ namespace URLNotesGrabberCORE
|
||||
Console.WriteLine($" Error deleting .bak files: {ex.Message}");
|
||||
}
|
||||
|
||||
RenameExistingTxtFilesToBak(ttFolderPath);
|
||||
RenameExistingTxtFilesToBak(folder);
|
||||
|
||||
var posts = DataAccess.GetAllPostsForBlog(blogName);
|
||||
Console.WriteLine($" Found {posts.Count} post(s) for this blog.");
|
||||
@@ -46,7 +69,7 @@ namespace URLNotesGrabberCORE
|
||||
foreach (var typeGroup in grouped)
|
||||
{
|
||||
string postType = typeGroup.Key ?? "Unknown";
|
||||
string outputFilePath = Path.Combine(ttFolderPath, $"{postType}.txt");
|
||||
string outputFilePath = Path.Combine(folder, $"{postType}.txt");
|
||||
var ordered = typeGroup.OrderBy(p => p.Date).ToList();
|
||||
Console.WriteLine($" Writing {ordered.Count} post(s) to {postType}.txt");
|
||||
|
||||
@@ -65,10 +88,57 @@ namespace URLNotesGrabberCORE
|
||||
}
|
||||
}
|
||||
|
||||
Console.WriteLine("\nOutput mode complete.");
|
||||
Console.WriteLine($"\nOutput mode complete. {writtenCount} blog(s) exported, {missingFolderCount} skipped for a missing folder.");
|
||||
|
||||
if (writtenCount == 0)
|
||||
Console.WriteLine("Every TTFolderPath points at a folder this machine cannot see. The paths were most likely written on another machine -- re-run --updatepaths <root> here so they match local drive letters.");
|
||||
|
||||
return 0;
|
||||
}
|
||||
|
||||
// Re-reads the TumblThree Index metadata into Blogs.TTFolderPath before exporting.
|
||||
// A TL.db synced between machines cannot hold one absolute path that is valid on
|
||||
// both, so the stored paths are only trustworthy on the machine that wrote them --
|
||||
// which makes this refresh part of a normal export rather than a separate chore.
|
||||
// Returns false only when the run should stop.
|
||||
private static bool RefreshPaths(IConfiguration config, string[] args)
|
||||
{
|
||||
var settings = config.GetSection("appSettings");
|
||||
|
||||
if (args.Any(a => string.Equals(a, "--norefresh", StringComparison.OrdinalIgnoreCase)))
|
||||
{
|
||||
Console.WriteLine("Path refresh skipped (--norefresh); exporting to whatever paths TL.db already holds.");
|
||||
return true;
|
||||
}
|
||||
|
||||
string? root = args.FirstOrDefault(a => !a.StartsWith("--", StringComparison.Ordinal))
|
||||
?? settings.GetValue<string>("PathTTRoot");
|
||||
|
||||
var result = UpdateBlogPathsRunner.Scan(root, verbose: false);
|
||||
|
||||
switch (result.Outcome)
|
||||
{
|
||||
case UpdateBlogPathsRunner.ScanOutcome.NoRootConfigured:
|
||||
Console.WriteLine("No TumblThree root configured (appSettings:PathTTRoot is empty and none was passed),");
|
||||
Console.WriteLine("so TTFolderPath was not refreshed. Pass one as --output <root> to refresh it.");
|
||||
return true;
|
||||
|
||||
case UpdateBlogPathsRunner.ScanOutcome.IndexFolderMissing:
|
||||
// Silently exporting stale paths here would defeat the point of folding
|
||||
// the refresh in, so a bad root is a hard stop.
|
||||
Console.WriteLine($"Index folder not found at: {result.IndexPath}");
|
||||
Console.WriteLine("Fix the root (or pass --norefresh to export the paths already in TL.db).");
|
||||
return false;
|
||||
|
||||
default:
|
||||
Console.WriteLine($"Refreshed paths from {result.IndexPath}: " +
|
||||
$"{result.MetadataFiles} metadata file(s), {result.Written} written, " +
|
||||
$"{result.Unchanged} already correct, {result.NoLocation} without a location, " +
|
||||
$"{result.NoMatchingRow} without a blog row, {result.Errors} error(s).");
|
||||
return true;
|
||||
}
|
||||
}
|
||||
|
||||
private static void RenameExistingTxtFilesToBak(string folderPath)
|
||||
{
|
||||
try
|
||||
|
||||
@@ -343,7 +343,7 @@ namespace URLNotesGrabberCORE
|
||||
break;
|
||||
|
||||
case "--output":
|
||||
exitCode = OutputMode.Run(config);
|
||||
exitCode = OutputMode.Run(config, args.Skip(1).ToArray());
|
||||
break;
|
||||
|
||||
case "--revert":
|
||||
@@ -439,7 +439,9 @@ namespace URLNotesGrabberCORE
|
||||
|
||||
Console.WriteLine("--ingest [blogname]\t Ingest Tumblr .txt exports from appSettings:PathTTRoot into TL.db (all blogs, or single blog if name given)");
|
||||
|
||||
Console.WriteLine("--output\t Export posts from TL.db back to .txt files in each blog's TTFolderPath");
|
||||
Console.WriteLine("--output [rootPath]\t Refresh Blogs.TTFolderPath from <root>\\Index (or appSettings:PathTTRoot), then export posts from TL.db back to .txt files in each blog's folder");
|
||||
|
||||
Console.WriteLine("--output --norefresh\t Export without refreshing TTFolderPath first");
|
||||
|
||||
Console.WriteLine("--revert [blogname]\t Recursively scan the PathInput tree and restore *.bak back to *.txt (current .txt saved as next-free .bkN); optional blogname filters by path substring");
|
||||
|
||||
|
||||
@@ -262,6 +262,47 @@ under its Posts and Notes pages. Removing a blog hides the blog, not what it col
|
||||
|
||||
---
|
||||
|
||||
## `Posts.IsActive` and `Notes.IsActive` — optional, and not in this database yet
|
||||
|
||||
The same flag is being extended to the two content tables, with the same meaning: `0` is
|
||||
removed, anything else — including `NULL` — is live. **Neither column exists in the live
|
||||
`TL.db` as of 2026-07-29**; the DDL quoted above for `Posts` and `Notes` is complete. Like
|
||||
`Blogs.IsActive`, they are written from outside this crawler.
|
||||
|
||||
The crawler therefore treats both as optional, and as nothing it owns:
|
||||
|
||||
- **It never writes them.** No `INSERT` column list names `IsActive`, no `UPDATE` sets it,
|
||||
and `MapPrefixToColumn` — the only place a column name is chosen at runtime — cannot map
|
||||
to it. Re-crawling a removed post or note refreshes its content and leaves the flag at
|
||||
`0`. There is no `INSERT OR REPLACE` on `Posts` or `Notes` for a default to be reset by.
|
||||
- **It filters on them only when they exist.** `HasIsActiveColumn` in `DataAccess.cs` asks
|
||||
`PRAGMA table_info` once per table per database path and caches the answer; the filter
|
||||
is `COALESCE(IsActive, 1) = 1`, and it is dropped entirely when the column is absent.
|
||||
Naming a missing column is a hard SQLite error, so this is what lets one build run
|
||||
against databases on both sides of the change. The cache lives for the process — adding
|
||||
the columns to a live database takes effect on the next run.
|
||||
|
||||
Every read that selects posts or notes carries the filter: `GetPosts`, `GetReplies`,
|
||||
`GetRepliesWithMissingText`, `GetRepliesWithFilledText`, `GetAllPostTextColumns`,
|
||||
`GetAllPostsForBlog`, `GetPost`, `GetPostByIdAnyBlog`, and the engagement queries that
|
||||
count or join `Notes` (`GetBlogs`, `GetBlogsAll`, `GetBlogsForLikes`). The one deliberate
|
||||
omission is the `LEFT JOIN Notes` in `GetPosts`: nothing is selected from it and it can
|
||||
neither add nor remove a row, so filtering it would buy nothing.
|
||||
|
||||
`LegacyPostsDbImporter` is unfiltered too — it reads a foreign legacy database whose
|
||||
`Posts` table is not this schema.
|
||||
|
||||
Two consequences worth stating plainly, both inherited from how `Blogs.IsActive` is
|
||||
handled:
|
||||
|
||||
1. **Removal hides a row; it does not freeze it.** The write paths are keyed on a post the
|
||||
caller already selected, so an ingest or a correction run still overwrites the content
|
||||
of a removed post. Only selection is filtered.
|
||||
2. **`NULL` is live.** Write `0` or `1`, not `NULL`, but a `NULL` leaves the row visible
|
||||
rather than stranding it.
|
||||
|
||||
---
|
||||
|
||||
## Reproducing the numbers
|
||||
|
||||
```sql
|
||||
|
||||
@@ -4,34 +4,64 @@ namespace URLNotesGrabberCORE
|
||||
{
|
||||
// Port of ThreeTxtFileHelper/UpdateBlogPaths.cs. Reads .tumblr / .tmblrpriv metadata
|
||||
// files from a root\Index folder and populates Blogs.TTFolderPath in TL.db.
|
||||
//
|
||||
// Scan() is the reusable engine: --updatepaths wraps it as a standalone command and
|
||||
// --output calls it as a refresh step, because a TL.db synced between machines cannot
|
||||
// hold one absolute path that is correct on both.
|
||||
public static class UpdateBlogPathsRunner
|
||||
{
|
||||
public static int Run(string rootPath)
|
||||
public enum ScanOutcome
|
||||
{
|
||||
Completed,
|
||||
NoRootConfigured,
|
||||
IndexFolderMissing
|
||||
}
|
||||
|
||||
public sealed class ScanResult
|
||||
{
|
||||
public ScanOutcome Outcome { get; init; }
|
||||
public string RootPath { get; init; } = string.Empty;
|
||||
public string IndexPath { get; init; } = string.Empty;
|
||||
public int MetadataFiles { get; init; }
|
||||
public int Written { get; init; }
|
||||
public int Unchanged { get; init; }
|
||||
public int NoLocation { get; init; }
|
||||
public int NoMatchingRow { get; init; }
|
||||
public int Errors { get; init; }
|
||||
}
|
||||
|
||||
// verbose: log a line per metadata file. --updatepaths wants that detail; --output
|
||||
// only wants the counts, since a few hundred lines before the export would bury it.
|
||||
public static ScanResult Scan(string? rootPath, bool verbose)
|
||||
{
|
||||
if (string.IsNullOrWhiteSpace(rootPath))
|
||||
{
|
||||
Console.WriteLine("UpdateBlogPaths: rootPath is required.");
|
||||
return 1;
|
||||
}
|
||||
return new ScanResult { Outcome = ScanOutcome.NoRootConfigured };
|
||||
|
||||
DataAccess.EnsureTTFileHelperColumnsExist();
|
||||
|
||||
string indexPath = Path.Combine(rootPath, "Index");
|
||||
if (!Directory.Exists(indexPath))
|
||||
{
|
||||
Console.WriteLine($"Index folder not found at: {indexPath}");
|
||||
return 1;
|
||||
return new ScanResult
|
||||
{
|
||||
Outcome = ScanOutcome.IndexFolderMissing,
|
||||
RootPath = rootPath,
|
||||
IndexPath = indexPath
|
||||
};
|
||||
}
|
||||
|
||||
Console.WriteLine($"Scanning Index folder: {indexPath}");
|
||||
|
||||
var blogFiles = Directory.GetFiles(indexPath, "*.tumblr")
|
||||
.Concat(Directory.GetFiles(indexPath, "*.tmblrpriv"))
|
||||
.ToList();
|
||||
|
||||
if (verbose)
|
||||
Console.WriteLine($"Found {blogFiles.Count} blog metadata files");
|
||||
|
||||
int updatedCount = 0;
|
||||
int unchangedCount = 0;
|
||||
int noLocationCount = 0;
|
||||
int noRowCount = 0;
|
||||
int errorCount = 0;
|
||||
|
||||
foreach (var blogFile in blogFiles)
|
||||
{
|
||||
@@ -44,27 +74,91 @@ namespace URLNotesGrabberCORE
|
||||
|
||||
if (root.TryGetProperty("FileDownloadLocation", out JsonElement locationElement))
|
||||
{
|
||||
string? fileDownloadLocation = locationElement.GetString();
|
||||
string? fileDownloadLocation = locationElement.GetString()?.Trim();
|
||||
if (!string.IsNullOrWhiteSpace(fileDownloadLocation))
|
||||
{
|
||||
DataAccess.SetBlogTTFolderPath(blogName, fileDownloadLocation);
|
||||
// Report the database's answer, not the fact that the file parsed.
|
||||
if (DataAccess.SetBlogTTFolderPath(blogName, fileDownloadLocation))
|
||||
{
|
||||
updatedCount++;
|
||||
if (verbose)
|
||||
Console.WriteLine($"Updated {blogName}: {fileDownloadLocation}");
|
||||
}
|
||||
else if (DataAccess.BlogExists(blogName))
|
||||
{
|
||||
unchangedCount++;
|
||||
}
|
||||
else
|
||||
{
|
||||
noRowCount++;
|
||||
Console.WriteLine($"No Blogs row named '{blogName}' -- path not stored (name may differ in case)");
|
||||
}
|
||||
}
|
||||
else
|
||||
{
|
||||
noLocationCount++;
|
||||
if (verbose)
|
||||
Console.WriteLine($"Empty FileDownloadLocation in {blogFile}");
|
||||
}
|
||||
}
|
||||
else
|
||||
{
|
||||
noLocationCount++;
|
||||
if (verbose)
|
||||
Console.WriteLine($"No FileDownloadLocation found in {blogFile}");
|
||||
}
|
||||
}
|
||||
catch (Exception ex)
|
||||
{
|
||||
errorCount++;
|
||||
Console.WriteLine($"Error processing {blogFile}: {ex.Message}");
|
||||
}
|
||||
}
|
||||
|
||||
Console.WriteLine($"\nUpdated {updatedCount} blogs with TTFolderPath");
|
||||
return 0;
|
||||
return new ScanResult
|
||||
{
|
||||
Outcome = ScanOutcome.Completed,
|
||||
RootPath = rootPath,
|
||||
IndexPath = indexPath,
|
||||
MetadataFiles = blogFiles.Count,
|
||||
Written = updatedCount,
|
||||
Unchanged = unchangedCount,
|
||||
NoLocation = noLocationCount,
|
||||
NoMatchingRow = noRowCount,
|
||||
Errors = errorCount
|
||||
};
|
||||
}
|
||||
|
||||
public static int Run(string rootPath)
|
||||
{
|
||||
if (string.IsNullOrWhiteSpace(rootPath))
|
||||
{
|
||||
Console.WriteLine("UpdateBlogPaths: rootPath is required.");
|
||||
return 1;
|
||||
}
|
||||
|
||||
string indexPath = Path.Combine(rootPath, "Index");
|
||||
Console.WriteLine($"Scanning Index folder: {indexPath}");
|
||||
|
||||
var result = Scan(rootPath, verbose: true);
|
||||
|
||||
if (result.Outcome == ScanOutcome.IndexFolderMissing)
|
||||
{
|
||||
Console.WriteLine($"Index folder not found at: {result.IndexPath}");
|
||||
return 1;
|
||||
}
|
||||
|
||||
Console.WriteLine($"\n========== UpdateBlogPaths summary ==========");
|
||||
Console.WriteLine($"Metadata files: {result.MetadataFiles}");
|
||||
Console.WriteLine($"TTFolderPath written: {result.Written}");
|
||||
Console.WriteLine($"Already correct: {result.Unchanged}");
|
||||
Console.WriteLine($"No FileDownloadLocation: {result.NoLocation}");
|
||||
Console.WriteLine($"No matching blog row: {result.NoMatchingRow}");
|
||||
Console.WriteLine($"Errors: {result.Errors}");
|
||||
|
||||
Console.WriteLine($"\nBlogs now holding a TTFolderPath: {DataAccess.CountBlogsWithTTFolderPath()}");
|
||||
|
||||
return result.Errors == 0 ? 0 : 2;
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
@@ -142,6 +142,9 @@ ORDER BY et.tbl;
|
||||
-- 1c. EXTRA / UNEXPECTED COLUMNS: present in the DB but not in the expected
|
||||
-- list above. Informational only -- e.g. a NEWER backup, or a column this
|
||||
-- script's expected-list hasn't been updated for. Not an error by itself.
|
||||
-- Posts.IsActive and Notes.IsActive are listed here and NOT in 1a on
|
||||
-- purpose: they are written by other tools, the app only reads them when
|
||||
-- present, and it must not be told to add them. See TL.db.md.
|
||||
WITH expected(tbl, col) AS (
|
||||
VALUES
|
||||
('Posts','BlogName'),('Posts','PostID'),('Posts','HasNotesGathered'),('Posts','reblogURL'),
|
||||
@@ -150,14 +153,14 @@ WITH expected(tbl, col) AS (
|
||||
('Posts','Quote'),('Posts','Body'),('Posts','Tags'),('Posts','Link'),('Posts','PhotoURL'),
|
||||
('Posts','PhotoCaption'),('Posts','DownloadedFiles'),('Posts','AudioCaption'),('Posts','Question'),
|
||||
('Posts','Answer'),('Posts','Title'),('Posts','ByLikes'),('Posts','RootBlogName'),('Posts','RootURL'),
|
||||
('Posts','DateModified'),('Posts','DateCreated'),('Posts','PostType'),
|
||||
('Posts','DateModified'),('Posts','DateCreated'),('Posts','PostType'),('Posts','IsActive'),
|
||||
('Blogs','BlogName'),('Blogs','HasBeenOutput'),('Blogs','IsActive'),('Blogs','DateAdded'),
|
||||
('Blogs','ByLikes'),('Blogs','DateModified'),('Blogs','DateCreated'),('Blogs','LikesPulled'),
|
||||
('Blogs','LikesCursor'),('Blogs','LikesNewestTimestamp'),('Blogs','LikesLastRefreshed'),
|
||||
('Blogs','LikesLastNewCount'),('Blogs','TTFolderPath'),
|
||||
('Notes','RootBlogName'),('Notes','PostID'),('Notes','NoteBlogName'),('Notes','TimeStamp'),
|
||||
('Notes','Type'),('Notes','DatetimeCrawled'),('Notes','DateModified'),('Notes','DateCreated'),
|
||||
('Notes','replyText'),
|
||||
('Notes','replyText'),('Notes','IsActive'),
|
||||
('DailyAPICount','Date'),('DailyAPICount','APICount'),
|
||||
('ApiKeyPoolState','KeyName'),('ApiKeyPoolState','RetryUntil'),
|
||||
('ApiKeyPoolMeta','Id'),('ApiKeyPoolMeta','LastIndex')
|
||||
|
||||
Reference in New Issue
Block a user