Directory / Feed fetchers
Feed fetchers 23
AmazonSellerInitiatedListing
Listed onlyAmazon · Feed fetchers
Amazon's seller-initiated fetcher, documented alongside AmazonProductDiscoverybot. It fetches a website URL that a seller has supplied so that Amazon can build a product page in its store from that page. Amazon's robots.txt statement on the page covers AmazonProductDiscoverybot only, so no robots behaviour is recorded for this agent.
Apple Podcasts (iTMS)
VerifiableApple · Feed fetchers
Apple's podcast feed fetcher, which crawls only URLs associated with content registered on Apple Podcasts. Apple documents that iTMS traffic may come from applebot.apple.com hosts and that it does not follow robots.txt because it is not a general search crawler.
Artemis Web Reader
Fully verifiableArtemis (jamesg.blog) · Feed fetchers
The feed poller behind Artemis, a hosted web reader that follows websites and blogs on behalf of its subscribers. It updates once per day and uses HEAD and conditional GET requests with If-Modified-Since and If-None-Match to reduce bandwidth. Artemis publishes the addresses it polls feeds from as a plain-text list.
BazQux Fetcher
Listed onlyBazQux Reader · Feed fetchers
The feed fetcher of the BazQux Reader hosted RSS service. It retrieves and periodically refreshes the RSS/Atom and comment feeds that users have subscribed to, typically no more than once an hour per feed. BazQux documents that the fetcher acts as an agent of those users and therefore ignores robots.txt.
Facebook Catalog
Listed onlyMeta · Feed fetchers
Meta's product-catalog fetcher, identified by the facebookcatalog user-agent. It retrieves merchant product-data feeds used to build and refresh commerce catalogs surfaced across Facebook and Instagram.
Fedicabot
Listed onlyFedica · Feed fetchers
The fetcher used by Fedica's social-media publishing platform. The operator documents that it does not crawl whole sites: it reads only the page or RSS feed a Fedica user has specified, downloading feed content so the user can schedule posts from it, and reading Open Graph metadata so a shared link renders as a card. Fedica publishes the robots.txt token but states no robots.txt policy and no address ranges.
Feedbin
VerifiableFeedbin · Feed fetchers
Feedbin's feed fetcher, which retrieves RSS/Atom feeds that users have subscribed to. Its user agent includes the internal feed id and current subscriber count, and Feedbin documents forward-confirmed reverse DNS in *.bot.feedbin.com as the way to verify its requests.
Feeder Crawler
Listed onlyFeeder · Feed fetchers
Crawler for Feeder, a hosted RSS reading service. The operator documents that it only fetches feeds users have actively subscribed to, that it acts as an agent on the user's behalf and therefore does not check robots.txt, and that rate limits can be arranged by contacting support. Its user agent is a browser-shaped string in which Feeder deliberately inserts the feeder.co identifier as the honest self-identification.
Feedly Fetcher
Listed onlyFeedly · Feed fetchers
Feedly's fetcher, which retrieves RSS/Atom feed URLs after a user has explicitly added them to their Feedly. Feedly documents that it behaves as a direct agent of the user rather than a robot, and does not publish a fixed IP list because its source IPs change over time.
Feedspot
Listed onlyFeedspot · Feed fetchers
Feedspot is a hosted content reader and feed aggregation service. Its bot fetches RSS and Atom feeds and web content on behalf of Feedspot users.
FeedWind Crawler
Listed onlyMikle KK (FeedWind) · Feed fetchers
The crawler behind FeedWind, an embeddable RSS/Atom widget. It fetches the feed sources that a widget owner has configured, every five minutes to five hours depending on their plan, so the rendered widget stays current. The operator's support page prints the crawler's full user agent and asks feed publishers to unblock it if a firewall is turning it away.
Flipboard Proxy
Listed onlyFlipboard, Inc. · Feed fetchers
Flipboard's proxy service, which fetches and prepares elements of a page (e.g. a social feed a user asked Flipboard to scan) for presentation in the Flipboard app. Flipboard's own docs say these requests currently originate from an Amazon EC2 cluster but publish no fixed IP list.
Feedfetcher-Google
Fully verifiableGoogle · Feed fetchers
Google's feed retrieval agent for RSS and Atom feeds used by Google News and WebSub. It fetches and periodically refreshes feeds that users of an app or service have explicitly subscribed to.
Hatena::Russia::Crawler
Listed onlyHatena Co., Ltd. (Hatelabo) · Feed fetchers
The fetcher behind Daichecker, the antenna service run on Hatelabo, Hatena's experimental-services lab. It checks the pages and feeds that users have registered for updates. Hatena documents that it parses only the robots.txt groups that name this user agent directly and does not apply the User-agent: * group, so a wildcard rule will not stop it. Hatena notes the name comes from an internal code name for RSS-reader development and has no connection to the country.
Innguma Fetcher
Listed onlyInnguma · Feed fetchers
Innguma's feed fetcher. It retrieves RSS and Atom feeds that users have added to Innguma or to another application built on the Innguma cloud, and refreshes them roughly once an hour. The operator states that because the requests follow explicit action by human users rather than automated crawling, the fetcher does not follow robots.txt, and documents serving an error status to the Innguma/1.0 user agent as the opt-out.
Inoreader Fetcher
Fully verifiableInnologica · Feed fetchers
Inoreader's feed fetcher, which retrieves RSS/Atom feeds that Inoreader users have subscribed to. Its own docs state it does not read robots.txt because it fetches specific, user-requested feed URLs rather than crawling a site, and it publishes a live list of its backend fetcher IPs.
Miniflux
Listed onlyMiniflux · Feed fetchers
Miniflux is a minimalist, open-source, self-hosted feed reader. User-run instances fetch the RSS and Atom feeds their subscribers add, identifying themselves with a Miniflux User-Agent.
NewsBlur Feed Fetcher
Listed onlyNewsBlur · Feed fetchers
NewsBlur's open-source feed fetcher, which polls RSS/Atom feeds on behalf of subscribed users. Its user agent embeds the live subscriber count and the feed's permalink; NewsBlur publishes no fixed IP range for it.
Productsup Website Crawler
Listed onlyProductsup · Feed fetchers
Crawler behind Productsup's website data-import source, which reads pages of a site a Productsup customer has configured and turns them into a product data feed. Productsup documents the default user agent, and notes that customers may append a hash to it for their own filtering.
rakutenusabot-image
Listed onlyRakuten · Feed fetchers
Rakuten's image-extraction bot, which retrieves product images from merchant sites that have partnered with Rakuten. Its bot page names the identifying token that appears in the user-agent string and gives an abuse contact.
ScourRSSBot
Listed onlyScour (Evan Schwartz) · Feed fetchers
Feed fetcher for Scour, a personalized content feed service in which users subscribe to RSS, Atom and JSON feeds and topics of interest. It polls each subscribed feed once every 900 seconds regardless of subscriber count, uses conditional requests, rate-limits itself to two requests per second per registrable domain, and backs off on error responses.
Superfeedr
VerifiableSuperfeedr · Feed fetchers
Superfeedr's PubSubHubbub feed-polling infrastructure, which fetches feed URLs that publishers or subscribers have registered with the service. Its docs publish a list of current node IPs but warn it changes as they add or remove cloud capacity.
W3C Feed Validation Service
VerifiableWorld Wide Web Consortium (W3C) · Feed fetchers
W3C's feed validator, which fetches a user-submitted RSS or Atom feed to check it against the feed specifications. It is a one-off conformance check rather than a subscription fetcher, and W3C publishes its user agent and the shared validator source addresses.