So, what I want is to download lemmy posts (including comments) as raw html through a script. The problem I have is, that curl by default does not include the css of the page. Is there a way I can download the posts webpages as they are being displayed?

  • macniel@feddit.org
    link
    fedilink
    arrow-up
    2
    ·
    2 days ago

    Mhm you could run a script through the HTML and look for stylesheet hrefs and download next to them.

    • da_cow (she/her)@feddit.orgOP
      link
      fedilink
      arrow-up
      2
      ·
      edit-2
      2 days ago

      I was more looking at doing this with a tool like wget. I got the following command, that downloads the post and its assets, but for some reason it then does not actually work:

      wget --recursive --no-clobber --page-requisites --html-extension --convert-links --restrict-file-names=windows --domains "feddit.org" --no-parent  "feddit.org/post" "https://feddit.org/post/33942692
      

      The advantage of wget is, that it also accepts multiple links at once, so it is quite easy do download multiple posts at once (which is my goal)

      • apparia@discuss.tchncs.de
        link
        fedilink
        English
        arrow-up
        4
        ·
        2 days ago

        If this is what you are seeing that makes it “not work”:

        It seems to be implemented in a script that runs when the saved page loads and hides the content, even though that content is actually there. You can tell wget to ignore scripts (which aren’t very useful in a download anyway, as they mostly pertain to interactivity) with --reject js. I tried this and was able to view the downloaded post properly.

        I’ve no idea why the script does that, one could probably look at the lemmy-ui source code to figure it out.