Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 7 additions & 3 deletions CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -274,19 +274,23 @@ Admins export the site from `/admin/exports` (`Admin::ExportsController`, admin

Downloads go through `Admin::ExportDownloadsController` (`GET /admin/exports/:id/download`), which uses `ActiveStorage::Streaming#send_blob_stream` behind `require_admin` rather than redirecting to an Active Storage blob URL — those signed URLs don't expire, and a JSON export contains every subscriber email. It's a separate controller because `ActiveStorage::Streaming` includes `ActionController::Live`. `rubyzip` must stay a direct `Gemfile` dependency: it's otherwise only pulled in transitively by `selenium-webdriver` in the test group.

### Content Import (WordPress, Ghost)
Admins upload an export at `/admin/imports` (`Admin::ImportsController`, admin role only; one `admin/imports/_upload_form` card per `Import.sources` key). `Import` (`source` enum `wordpress`/`ghost`, `validate: true`; `status` pending/processing/completed/failed; `has_one_attached :file` ≤ `Import::MAX_FILE_SIZE` (50 MB) whose type must match `Import::ACCEPTED_FILES[source]` (XML for WordPress, JSON for Ghost; error keys `invalid_<source>_type`); optional `site_url`, normalized and validated as http(s); `stats` JSON) is processed by `ImportJob`, which opens the stored file, runs `Import#importer` (`Import.importer_for(source)`) with `io:`, `user:`, `site_url:`, saves the returned stats, and broadcasts the `admin/imports/_import` partial on the `"imports"` stream. To add a source (e.g. Substack): add the enum value, an `ACCEPTED_FILES` entry, a branch in `Import.importer_for`, locale keys (`admin.imports.sources.*`, `admin.imports.index.<source>_description`/`_file_label`), and an importer.
### Content Import (WordPress, Ghost, Substack)
Admins upload an export at `/admin/imports` (`Admin::ImportsController`, admin role only; one `admin/imports/_upload_form` card per `Import.sources` key). `Import` (`source` enum `wordpress`/`ghost`/`substack`, `validate: true`; `status` pending/processing/completed/failed; `has_one_attached :file` ≤ `Import::MAX_FILE_SIZE` (50 MB) whose content type or extension must match `Import::ACCEPTED_FILES[source]` (XML for WordPress, JSON for Ghost, zip or CSV for Substack; error keys `invalid_<source>_type`); optional `site_url`, normalized and validated as http(s); `stats` JSON) is processed by `ImportJob`, which opens the stored file, runs `Import#importer` (`Import.importer_for(source)`) with `io:`, `user:`, `site_url:`, saves the returned stats, and broadcasts the `admin/imports/_import` partial on the `"imports"` stream. To add a source (e.g. Ghost members CSV): add the enum value, an `ACCEPTED_FILES` entry, a branch in `Import.importer_for`, locale keys (`admin.imports.sources.*`, `admin.imports.index.<source>_description`/`_file_label`), and an importer.

**Importers** subclass `Imports::BaseImporter` and implement `import_items`; `#call` → stats hash (`posts_imported`, `pages_imported`, `skipped`, `failed`, `media_downloaded`, `media_failed`, `warnings` capped at `Import::MAX_WARNINGS`, `warnings_truncated`). The base provides `importable_slug?` (skips existing slugs and `Page::RESERVED_SLUGS`, making re-runs idempotent), `import_item` (confines an unexpected `StandardError` — odd markup crashing a converter, a storage failure — to that item: logged, recorded as a warning, counted in `failed`, import continues; failures reading the export happen outside it and still fail the whole import), `save_item` (per-item transaction; `RecordInvalid` becomes a warning + `failed`), `attach_featured_image`, `find_or_create_category`/`find_or_create_tag` (match slug, then name), `resolve_schedule` (scheduled items whose date passed become published), and purging downloaded blobs left unattached. All content is assigned to the importing admin. **Each item is saved exactly once with content/featured image/taxonomy already set** — that is what keeps imports silent: notifications only fire from `Post#publish!`, and `Post::Webhookable`/`Versionable` only react to updates. Don't split an imported post into create-then-update. Media is downloaded *before* the per-item transaction so SQLite's write lock isn't held across HTTP.

- `Imports::WordpressImporter`: `publish`→published, `future`→scheduled, `draft`/`pending`/`private`→draft, anything else skipped. First non-"uncategorized" category → `category`; remaining categories + tags → `tags`. Excerpt → `meta_description`. `Imports::Wordpress::WxrParser` reads namespace URIs from the document (WXR 1.0–1.2 differ; `remove_namespaces!` would collapse `content:encoded`/`excerpt:encoded`), parses with `nonet` (no external entities), and resolves `_thumbnail_id` postmeta to attachment URLs.
- `Imports::GhostImporter`: `published`→published, `scheduled`→scheduled, `draft` and email-only `sent`→draft (with a warning). Visibility `public`/`members`/`paid`+`tiers` → `public`/`members_only`/`paid_only`. `featured` kept, `custom_excerpt` → `subtitle`, `meta_description` (post or `posts_meta`) → `meta_description`. The primary (first public, by `posts_tags.sort_order`) tag → `category`, all public tags → `tags`; internal `#tags` are dropped. `Imports::Ghost::ExportParser` accepts both `{"db":[{meta,data}]}` and bare `{meta,data}`, forces UTF-8 (the job opens files in binary mode), and picks content from `html`, else renders `mobiledoc` via `Imports::Ghost::MobiledocRenderer`, else escapes `plaintext` into paragraphs (Lexical-only posts; reported as a warning).

- `Imports::SubstackImporter`: posts come from `posts.csv` + `posts/<post_id>.html` (slug is the part of `post_id` after the first dot); `is_published` → published/scheduled by `post_date`, else draft; audience `everyone`/`only_free`/`only_paid`+`founding` → `public`/`members_only`/`paid_only`; threads are skipped, podcasts import their show notes with a warning. **Subscribers** from `email_list*.csv` are created with `Subscriber.new` — never `subscribe_or_sign_in!`, which would send confirmation mail and fire `subscriber.created` webhooks — as confirmed with `confirmed_at`/`created_at` set to the Substack signup date; `email_disabled` rows become unsubscribed; existing (or invalid) emails are skipped and counted in `subscribers_imported`/`subscribers_skipped`. Paid billing can't move off Substack's Stripe account, so `active_subscription` rows get a "Substack paid" or (plan `comp`/`gift`) "Substack comp" `SubscriberLabel` instead of a membership. Subscriber rows are saved in transactions of `SUBSCRIBER_BATCH_SIZE` (500) because a commit per row is slow on SQLite, each inside its own savepoint (`transaction(requires_new: true)`) so one failing row rolls back alone rather than taking the batch with it; labels are resolved before the savepoint so a rolled-back row can't invalidate a memoized label. `Imports::Substack::ExportReader` detects zip vs. bare CSV by magic bytes, matches zip entries by path suffix (so a wrapping folder works), never extracts to disk, and enforces `MAX_ENTRIES`/`MAX_ENTRY_BYTES`/`MAX_TOTAL_BYTES` while decompressing (declared entry sizes can lie). CSVs are parsed with the `csv` gem — a direct `Gemfile` dependency because Ruby 3.4 no longer ships it as a default gem — with the BOM stripped and headers downcased.

**Converters** subclass `Imports::ContentConverter`, whose `convert(html, title:)` pipeline is: `preprocess` hook → drop unsafe tags → YouTube/Vimeo iframes become links (other iframes removed with a warning) → `transform(fragment)` hook → each `<img>` is downloaded into an `ActionText::Attachment` (candidates from the overridable `image_candidates`; `<figcaption>` becomes the caption; a failed download keeps the remote `<img>`) → strip `class`/`style`/`srcset`/`data-*`/`on*` etc. → drop empty paragraphs. Attributes are stripped *after* `transform` so subclasses can match source class names. Image containers are resolved for all images before any is replaced (otherwise replacing one gallery image makes its sibling look like the figure's only image and deletes the figure).
- `Imports::Wordpress::ContentConverter`: strips Gutenberg `<!-- wp:* -->` comments, runs a `wpautop` port only on non-block (classic) content, converts `[caption]` → figure, `[embed]`/`[youtube]` → links, removes `[gallery]`/`[audio]`/`[video]`/`[playlist]` with a warning, **leaves unknown shortcodes as text** (with a warning) so bracketed prose isn't lost, turns Gutenberg embed figures into links, and prefers a linked full-size image over the resized `src`.
- `Imports::Ghost::ContentConverter`: replaces `__GHOST_URL__` with the import's `site_url` (Ghost exports contain no domain; without it those images fail to download and the importer warns once), strips `<!--kg-card-begin/end-->` markers, maps Koenig cards (bookmark/button/file → link, callout → blockquote, audio/video removed with a warning, signup removed, other `kg-*` wrappers unwrapped), and tries the original image before `/content/images/size/wNNN/` variants.

`Imports::MediaDownloader` is the SSRF-safe fetcher: every URL and redirect hop (max 3) goes through `Webhooks::UrlGuard.resolve!` with the connection pinned via `Net::HTTP#ipaddr=`, bodies are capped at 10 MB while streaming, both the header and the sniffed (Marcel) type must be `image/*`, results are memoized per URL, and it never raises (failures are recorded in `#failures`). Tests stub `Net::HTTP.new` and `Webhooks::UrlGuard.resolve` or inject a fake downloader rather than hitting the network; fixture exports are `test/fixtures/files/wordpress.xml` and `test/fixtures/files/ghost.json`.
- `Imports::Substack::ContentConverter`: reads embed JSON from `data-attrs` (the base strips `data-*` only after `transform`): subscribe widgets, paywall markers, SVG icons and subscribe/share buttons are removed, other buttons become links, `youtube-wrap`/`vimeo-wrap`/`tweet`/`embedded-post-wrap`/Spotify/SoundCloud become links, `pullquote` becomes a blockquote, and `picture`/image insets are unwrapped. Images try the original upload encoded at the end of `substackcdn.com/image/fetch/<transforms>/<url>` first.

`Imports::MediaDownloader` is the SSRF-safe fetcher: every URL and redirect hop (max 3) goes through `Webhooks::UrlGuard.resolve!` with the connection pinned via `Net::HTTP#ipaddr=`, bodies are capped at 10 MB while streaming, both the header and the sniffed (Marcel) type must be `image/*`, results are memoized per URL, and it never raises (failures are recorded in `#failures`). Tests stub `Net::HTTP.new` and `Webhooks::UrlGuard.resolve` or inject a fake downloader rather than hitting the network; fixture exports are `test/fixtures/files/wordpress.xml`, `test/fixtures/files/ghost.json`, and the `test/fixtures/files/substack/` directory, which `SubstackExportHelper#substack_export_zip` (`test/test_helpers/substack_export_helper.rb`, loaded with `require_relative`) zips in memory.

### Authentication
- **Admin**: session-based (signed cookie, 14-day expiry). The `Authentication` concern owns cookie → `Session` resumption (`resume_session`, memoized via a `Current.session` short-circuit) and is included once on `ApplicationController`, so both admin (`current_user`) and identity (`IdentityAuthentication#current_identity`) lookups share a single `sessions` query per request.
Expand Down
3 changes: 3 additions & 0 deletions Gemfile
Original file line number Diff line number Diff line change
Expand Up @@ -68,6 +68,9 @@ gem "commonmarker", "~> 2.3"
# HTML to Markdown conversion for content export [https://github.com/xijo/reverse_markdown]
gem "reverse_markdown", "~> 3.0"

# CSV parsing for Substack imports (no longer a default gem as of Ruby 3.4)
gem "csv", "~> 3.3"

# Zip archives for Markdown content export [https://github.com/rubyzip/rubyzip]
gem "rubyzip", "~> 3.0", require: "zip"

Expand Down
3 changes: 3 additions & 0 deletions Gemfile.lock
Original file line number Diff line number Diff line change
Expand Up @@ -117,6 +117,7 @@ GEM
cbor (~> 0.5.9)
openssl-signature_algorithm (~> 1.0)
crass (1.0.7)
csv (3.3.6)
date (3.5.1)
debug (1.11.1)
irb (~> 1.10)
Expand Down Expand Up @@ -484,6 +485,7 @@ DEPENDENCIES
bundler-audit
capybara
commonmarker (~> 2.3)
csv (~> 3.3)
debug
diffy
image_processing (~> 2.1)
Expand Down Expand Up @@ -555,6 +557,7 @@ CHECKSUMS
connection_pool (3.0.2) sha256=33fff5ba71a12d2aa26cb72b1db8bba2a1a01823559fb01d29eb74c286e62e0a
cose (1.3.1) sha256=d5d4dbcd6b035d513edc4e1ab9bc10e9ce13b4011c96e3d1b8fe5e6413fd6de5
crass (1.0.7) sha256=94868719948664c89ddcaf0a37c65048413dfcb1c869470a5f7a7ceb5390b295
csv (3.3.6) sha256=aba61e7e507a66f03d45cb1f3c4b6359861c3504038b422962875dce099e4456
date (3.5.1) sha256=750d06384d7b9c15d562c76291407d89e368dda4d4fff957eb94962d325a0dc0
debug (1.11.1) sha256=2e0b0ac6119f2207a6f8ac7d4a73ca8eb4e440f64da0a3136c30343146e952b6
diffy (3.4.4) sha256=79384ab5ca82d0e115b2771f0961e27c164c456074bd2ec46b637ebf7b6e47e3
Expand Down
1 change: 1 addition & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -12,6 +12,7 @@ A self-hosted blogging platform built with Ruby on Rails 8.1 and the Solid stack
- **Content Export** — Download your whole site as a Markdown zip (one file per post/page with YAML front matter, plus images) or a full JSON backup including subscribers; exports run in the background from **Admin → Import & Export**
- **WordPress Import** — Upload a WordPress export (WXR `.xml`) to bring over posts, pages, categories, tags, and images; shortcodes and embeds are converted, and re-running an import skips anything already imported
- **Ghost Import** — Upload a Ghost JSON export to bring over posts, pages, tags, featured posts, and member/paid visibility; Koenig cards and Mobiledoc content are converted, and images are downloaded when you provide your Ghost site URL
- **Substack Import** — Upload your Substack export .zip to bring over posts (with paid/free audiences) and your subscriber list; subscribers are added as confirmed without sending any emails, and paid/comp subscribers are labelled for follow-up
- **Content Organization** — Categories, tags with searchable combo box and inline creation
- **Reader Engagement** — Comments with threading and moderation, loves, social share buttons, subscriber magic-link auth, email notifications
- **Social Embeds** — X/Twitter and YouTube via oEmbed
Expand Down
11 changes: 7 additions & 4 deletions app/models/import.rb
Original file line number Diff line number Diff line change
@@ -1,12 +1,13 @@
class Import < ApplicationRecord
MAX_FILE_SIZE = 50.megabytes
ACCEPTED_FILES = {
"wordpress" => { content_types: %w[application/xml text/xml application/rss+xml], extension: "xml" },
"ghost" => { content_types: %w[application/json], extension: "json" }
"wordpress" => { content_types: %w[application/xml text/xml application/rss+xml], extensions: %w[xml] },
"ghost" => { content_types: %w[application/json], extensions: %w[json] },
"substack" => { content_types: %w[application/zip application/x-zip-compressed text/csv], extensions: %w[zip csv] }
}.freeze
MAX_WARNINGS = 50

enum :source, { wordpress: 0, ghost: 1 }, prefix: :source, validate: true
enum :source, { wordpress: 0, ghost: 1, substack: 2 }, prefix: :source, validate: true
enum :status, { pending: 0, processing: 1, completed: 2, failed: 3 }

belongs_to :user
Expand All @@ -23,6 +24,7 @@ def self.importer_for(source)
case source.to_s
when "wordpress" then Imports::WordpressImporter
when "ghost" then Imports::GhostImporter
when "substack" then Imports::SubstackImporter
else raise ArgumentError, "Unknown import source: #{source}"
end
end
Expand Down Expand Up @@ -65,7 +67,8 @@ def file_is_acceptable

def accepted_file_type?
accepted = ACCEPTED_FILES.fetch(source)
accepted[:content_types].include?(file.blob.content_type) || file.blob.filename.extension.casecmp?(accepted[:extension])
extension = file.blob.filename.extension.to_s.downcase
accepted[:content_types].include?(file.blob.content_type) || accepted[:extensions].include?(extension)
end

def site_url_is_http
Expand Down
103 changes: 103 additions & 0 deletions app/services/imports/substack/content_converter.rb
Original file line number Diff line number Diff line change
@@ -0,0 +1,103 @@
module Imports
module Substack
# Substack-specific conversion on top of Imports::ContentConverter.
#
# Substack post bodies are editor output with embeds serialized as JSON in
# data-attrs (base strips data-* only after #transform, so it's readable):
#
# subscribe widgets / paywall markers / subscribe & share buttons → removed
# other buttons → link paragraph
# youtube-wrap / vimeo-wrap → video link
# tweet / embedded-post-wrap / spotify / soundcloud → link paragraph
# pullquote → blockquote
# captioned-image-container / picture / image insets → unwrapped
#
# Images are served through substackcdn.com/image/fetch/<transforms>/<url>;
# the original upload is the URL-encoded last segment, so it is tried first.
class ContentConverter < Imports::ContentConverter
REMOVED = [
".subscription-widget-wrap", ".subscription-widget-wrap-editor", ".subscription-widget",
"[data-component-name^='SubscribeWidget']", ".paywall-jump", "[data-component-name^='Paywall']",
".image-link-expand", "svg", "picture source"
].freeze
LINK_EMBEDS = {
".tweet" => "url", ".embedded-post-wrap" => "url", ".spotify-wrap" => "url",
".soundcloud-wrap" => "url", ".apple-podcast-container" => "url"
}.freeze
SUBSCRIBE_OR_SHARE = %r{/subscribe\b|[?&]action=share|/share\b|substack\.com/refer}i
CDN_FETCH = %r{\Ahttps?://substackcdn\.com/image/fetch/[^/]+/(.+)\z}i

private

def transform(fragment)
fragment.css(REMOVED.join(",")).each(&:remove)
convert_buttons(fragment)
convert_videos(fragment)
LINK_EMBEDS.each do |selector, key|
fragment.css(selector).each do |node|
attrs = data_attrs(node)
replace_with_link(node, attrs[key], attrs["title"])
end
end
fragment.css(".pullquote").each { |node| node.name = "blockquote" }
fragment.css("picture, .image2-inset, .captioned-image-container").reverse_each { |node| unwrap(node) }
end

def convert_buttons(fragment)
fragment.css(".button-wrapper, .captioned-button-wrap").each do |wrapper|
anchor = wrapper.at_css("a")
href = anchor&.[]("href").to_s
if href.blank? || href.match?(SUBSCRIBE_OR_SHARE)
wrapper.remove
else
replace_with_link(wrapper, href, anchor.text)
end
end
end

def convert_videos(fragment)
fragment.css(".youtube-wrap").each do |node|
video_id = data_attrs(node)["videoId"]
replace_with_link(node, video_id.present? ? "https://www.youtube.com/watch?v=#{video_id}" : nil)
end
fragment.css(".vimeo-wrap").each do |node|
video_id = data_attrs(node)["videoId"]
replace_with_link(node, video_id.present? ? "https://vimeo.com/#{video_id}" : nil)
end
end

def image_candidates(img)
href = img.ancestors("a").first&.[]("href").to_s
stored = data_attrs(img)["src"]
[ original_url(href), stored, original_url(img["src"].to_s), img["src"] ]
.compact_blank.select { |url| url.match?(%r{\Ahttps?://}i) }
end

def original_url(url)
encoded = url[CDN_FETCH, 1]
return (url.match?(IMAGE_EXTENSION) ? url : nil) unless encoded

CGI.unescape(encoded)
end

def data_attrs(node)
JSON.parse(node["data-attrs"].to_s)
rescue JSON::ParserError
{}
end

def replace_with_link(node, url, text = nil)
if url.to_s.match?(%r{\Ahttps?://}i)
node.replace("<p>#{link(url, text.to_s.squish.presence || url)}</p>")
else
node.remove
end
end

def unwrap(node)
node.add_next_sibling(node.children)
node.remove
end
end
end
end
Loading