I got tired of checking 8 sites for AI news so I wrote a scraper that does it for me
Post original
Every morning I'd open HN, Reddit, GitHub Trending, Hugging Face papers, Ars Technica... just to see if I missed anything interesting in AI. Half the tabs were the same story repeated. So I wrote a pipeline that scrapes all of them, deduplicates by URL + semantic similarity, feeds the top 9 into a free LLM (Gemma 4 via OpenRouter) to write short summaries, and builds a static site. The trickiest part was the dedup: same paper hits HN, Reddit, and Hugging Face under slightly different titles. Simhash + URL dedup works well enough. It's been running for a few weeks now and I actually use it daily. Total cost: $0. It's not perfect and I'm still tuning it for content filtering, but I find it useful.   submitted by   /u/TurboBanano [link]   [comments]
Rascunhos
Sem rascunho (score abaixo do threshold). Ajuste o threshold em Configurações se quiser gerar rascunho para leads com score menor.