So, what I want is to download lemmy posts (including comments) as raw html through a script. The problem I have is, that curl by default does not include the css of the page. Is there a way I can download the posts webpages as they are being displayed?
Fwiw there’s SinglePage extension for web browsers. Maybe you can get a headless way to run that… Which frontend would you save from? I’d say i rather prefer to get json data from backend
I am preferring the standard Lemmy UI, but I already got it working.
why html? just saving the json it fetches in the first place would be a lot easier to work with IMO, but it depends on what you’re doing with it
I second that - can’t you just use the API and request the post and its comments as json?
Unless you want to explicitly archive the post as a human viewable HTML, the structured json from the API is almost always superior to further analyze, archive,… the posts.
Archiving thwmin a human readable form is exactly what I want to do.
Then I’m not sure if curl is up to the task in this case - most modern websites are PWAs that dynamically request data from their Backen and display it. Curl just fetches the initial HTML but doesn’t execute the JS and thus might not see the content of the page as it isn’t there yet the moment it gets saved.
Some PWAs “render” the first page contents in your HTML so that it is included and do not need to be requested.
Maybe try the alternative frontends if some of them work better? Also remember that feddit.org and many other instances deploy Anubis and might lock you out from making simple requests with curl
Got it working with wget
Mhm you could run a script through the HTML and look for stylesheet hrefs and download next to them.
I was more looking at doing this with a tool like wget. I got the following command, that downloads the post and its assets, but for some reason it then does not actually work:
wget --recursive --no-clobber --page-requisites --html-extension --convert-links --restrict-file-names=windows --domains "feddit.org" --no-parent "feddit.org/post" "https://feddit.org/post/33942692The advantage of wget is, that it also accepts multiple links at once, so it is quite easy do download multiple posts at once (which is my goal)
If this is what you are seeing that makes it “not work”:

It seems to be implemented in a script that runs when the saved page loads and hides the content, even though that content is actually there. You can tell wget to ignore scripts (which aren’t very useful in a download anyway, as they mostly pertain to interactivity) with
--reject js. I tried this and was able to view the downloaded post properly.I’ve no idea why the script does that, one could probably look at the
lemmy-uisource code to figure it out.That was my exact problem. Thanks for the solution.
Your browser dev tools network tab can be set to filter by CSS files. Then you just copy those URLs, and sort out the unnecessary garbage if there is some (ad frameworks, etc.). Since CSS won’t really be changing a lot, you can just download it once and include it statically using your script and some
<link rel="stylesheet" href="./path/to/file.css">tags in the<head>





