<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Data Mining on SoloSoft</title><link>https://www.solosoft.dev/tags/data-mining/</link><description>Recent content in Data Mining on SoloSoft</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Fri, 01 May 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://www.solosoft.dev/tags/data-mining/index.xml" rel="self" type="application/rss+xml"/><item><title>MediaCrawler: Open-Source Social Media Data Scraper with 30K Stars</title><link>https://www.solosoft.dev/post/mediacrawler-social-scraper-2026/</link><pubDate>Fri, 01 May 2026 00:00:00 +0000</pubDate><guid>https://www.solosoft.dev/post/mediacrawler-social-scraper-2026/</guid><description>&lt;p&gt;Social media data is a goldmine for market research, trend analysis, and competitive intelligence &amp;ndash; but accessing it programmatically is notoriously difficult. Platforms actively block scrapers, change their APIs, and require complex authentication flows. &lt;strong&gt;MediaCrawler&lt;/strong&gt; has emerged as one of the most popular open-source solutions to this challenge, with over 30,000 GitHub stars and support for all major Chinese social media platforms.&lt;/p&gt;
&lt;p&gt;The project at &lt;a href="https://github.com/NanmiCoder/MediaCrawler"&gt;github.com/NanmiCoder/MediaCrawler&lt;/a&gt; provides a unified framework for crawling data from Xiaohongshu (Little Red Book), Douyin (TikTok China), Kuaishou, Bilibili, Weibo, and more. It uses Playwright for browser automation, IP rotation, and cookie management to bypass anti-scraping measures. The result is a reliable data pipeline for extracting posts, comments, user profiles, and engagement metrics.&lt;/p&gt;</description></item></channel></rss>