Most mass scrapers, on the other hand, simply grab the raw HTML underneath. ShieldFont exploits this difference through an automated process called OpenType glyph substitution.

That said, because the whole defense rests on scrapers reading code rather than screens, taking a screenshot of a shielded page and running OCR on the image can still recover the real words.

Screen readers used by blind readers also work from the code, so they read the decoys aloud. ShieldFont ships with a beta feature that provides those readers with the real text instead.

    • T156@lemmy.world
      link
      fedilink
      English
      arrow-up
      2
      ·
      1 hour ago

      Presumably the AI scraper would also have OCR, and would sidestep things like this?

      • sudo@programming.dev
        link
        fedilink
        English
        arrow-up
        2
        ·
        29 minutes ago

        A scraper has many more ways around something like sheildfont than just OCR. The question will be if it was actually programmed to check for such measures.