

Discussion summary
Pulpie is a new model designed for cleaning web content, with users discussing its effectiveness, especially on JavaScript-heavy pages. Some suggest using headless browsers or existing tools like Trafilatura for similar tasks.
What the discussion says
- Users find Pulpie interesting and useful for web content extraction.
- Some recommend traditional tools like HTML to Markdown converters.
- Concerns about handling complex pages with JavaScript or diverse web structures.
“You'd typically use a headless browser to generate the fully rendered page.”
“It still has trouble with some pages, especially complex ones.”
Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- Very nice! Thank you for building this.by lnenad
- Why not use a plain old html → markdown converter? You can easily strip out ads using CSS /jQuery-like selectors. That would cost zero dollars.
- This model looks very interesting. How does it compare to heuristic-based approaches like Defuddle and Readability?by goldenjm
- Does it work with ecommerce for product scraping? E.g. Amazon, or Shopee (big in SEA)by wiradikusuma
- It's good looking, and I liked it. The trial page accessed from the hugging face website is a very inefficient experience when I use Mozilla and the dark theme, FYI.by kocamaz
- So this is tailored towards kind of a "reader view" for models right? Can it handle images, tables, shadow DOMs too? Like there are 3 use cases I have now - one is a simple text view for models to understand it, one is a "web clip" mode which would ideally preserve images and media, and one is to extract tabular data from web pages. Which ones is this good at?
- Why does the 'Quality vs Cost of Web Content Extraction' chart not have zero cost at the origin? Up to the right does not have to mean better; we can read.by esafak
- I did some research on this about 10 years ago. I spent 2 days hand labelling data from scraped news sites. Then built a good old fashioned Random Forest model to classify html nodes based on some feature engineering. turns out the P tag and the number-of-words threshold get you 90% of the way there, on news sites anyway. Great thing about RF models is they tell you which features are the most important. fun little project (apart from the 2 days of data labelling).by cpill