Draft for reviewNot yet reviewed, and not yet in effect.
Data sources
Data sources and takedown
1stSeen builds its forecasts only from public job postings and public archives of career pages. This page says what it reads, how it reads it, and how a company can ask to be left out.
Last updated
What 1stSeen reads
- Job-board APIs. The public endpoints that Greenhouse, Lever, Ashby, and SmartRecruiters publish for listing a company's open jobs.
- Company career pages, and the sitemaps and job feeds those sites publish. Collection reads only recruiting pages, such as a careers, jobs, or students page and the job pages a sitemap lists, never a blog, news, or product page. When a company is added, its home page, robots.txt, and sitemaps are read to find them.
- The Internet Archive's Wayback Machine. Its index and its archived copies of a company's careers or campus page, over the last five years, at most 45 copies per page by default. This is how past openings are dated.
- Reddit is off by default. It can be turned on only through Reddit's approved API, for an allowlist of communities. No author is recorded, and a post can only support a forecast, never date an opening.
Apart from Reddit's API when it is on, 1stSeen uses no account to read a source, and reads nothing behind a sign-in. Methodology explains how what it reads becomes a forecast.
What it keeps
For each posting or archived page: its address, when it was seen, its title and location, a publication date if the source states one, and at most 64 KB of its text, with a fingerprint of the content. For a career page watched for changes, at most 64 KB of its visible text, to tell what was added. Page code is not stored. Every forecast can show the evidence it was made from.
How often, and how politely
- On the current schedule, current postings are read four times a day, career pages for changes four times a day, and the archive once a week.
- Requests to any one site are spaced out: by default at least a quarter of a second apart for current postings, and a second and a half apart for career pages and the archive.
- By default each request gives up after 20 seconds and reads at most 10 MB, and a sitemap is followed at most two levels deep, five sitemap files and 25 job pages per run.
- A failed request is not retried until the next scheduled run. When a site answers with an access check or a challenge page, collection stops there. It uses no proxies and nothing that gets around a block.
- Every request identifies itself as 1stSeenEvidenceBot/0.2.
robots.txt
Before it requests anything from a site, collection reads that site's robots.txt, once per site per run, and skips every address it disallows for 1stSeenEvidenceBot or for all crawlers. The job-board APIs and the Wayback Machine are treated the same way.
- If a site has no robots.txt, nothing is skipped.
- If its robots.txt cannot be read because the server fails or does not answer, the whole site is skipped for that run.
- Each skip is recorded with the rule that caused it.
Company names and trademarks
Company and program names are used only to say which company and which program a forecast is about. They remain their owners' trademarks. 1stSeen shows no company logos. It is independent: no company it lists is affiliated with it or has endorsed it.
Asking for a company to be removed
If you speak for a company and want 1stSeen to stop reading its sites, or to take the company off 1stSeen, write to hello@1stseen.win and name the company and its website.
- You get a reply within two business days. It may ask you to confirm the request from an address at the company's own domain.
- Once the request is confirmed, and within five business days, collection from the sites you name stops, before the next scheduled run. If you asked for the company to be removed, its programs also leave every page, the watchlists that followed them, and the agent's answers, and every other program's forecast that used its postings is recomputed without them.
- You get a reply saying what was done, and when.
What was already collected is kept, out of sight of every page, unless you also ask for it to be erased. Then it is erased within 30 days, and 1stSeen keeps only the company's name, its web addresses, and the record of your request, so that it is never collected again.