Proxyrack - October 7, 2026

Scrape TikTok Search Results

Data ScrapingTutorials

A beginner-friendly guide for Linux. No programming knowledge needed - just follow each step in order. Written for a fresh server (or a normal desktop) with nothing installed.


Read this first. This guide assumes your computer has nothing installed for this project - no unzip, no curl, no uv, and no browser. Every tool gets its own install step below. Do the steps in order and do not skip any. The five most common mistakes are:

  1. cd tiktok fails - because the zip has no tiktok folder inside it. Step 3 fixes this.

  2. uv sync fails - because uv is not installed by default. Step 4 fixes this.

  3. Trying to install Python: you don't need to. uv supplies Python itself. The python3.12 package does not exist on many Ubuntu/Debian versions.

  4. Skipping xvfb: you must install it. Without xvfb-run headless fetching is impossible. Step 5 fixes this.

  5. Running headless first in Step 8 - wrong order. You must run Step 8A headed-under-Xvfb first (it will fail with 0 videos on purpose — ~90% of the time), and only then run Step 8B headless (retry twice for ~11/12 videos).


What We're Building

A small program that:

  1. Opens a real browser (Brave or Chrome) and visits a TikTok search results page

  2. Scrolls down so more videos load

  3. Saves the page HTML locally

  4. Pulls out each video (URL, author, description, likes, hashtags, thumbnail) and prints it

TikTok builds its search page with JavaScript, so the raw page has no results in it. We use a real browser controlled through Chrome DevTools Protocol (CDP). The browser loads the page, scrolls, and grabs the finished HTML. After that first fetch, everything runs offline against the saved file.

The browser runs headless (invisible) for the final fetch, but headless cannot work on its own. You must first do one headed init run under a virtual screen called Xvfb (Step 5 + Step 8A). That headed run invokes TikTok's JavaScript to verify your session and cookies — it fails with 0 videos ~90% of the time on purpose. Only after it fails do you run the headless fetch (Step 8B), which succeeds on the second try with about 11/12 videos.

You should have received one file: tiktok.zip. It already contains the whole program - including the parser, the API, and the command-line tool. You do not copy code from any other project and you do not type the code by hand.


Step 1: Open a Terminal

What is a terminal?

A terminal is a window where you type commands instead of clicking buttons. Every command in this guide is typed here, then you press Enter.

If you are on a server (a droplet/VPS), you are already inside a terminal after you connect. Skip to the check below.

To open it on a Linux desktop:

  1. Press Ctrl + Alt + T at the same time, or

  2. Open your applications menu and search for Terminal, Console, or GNOME Terminal

Check it works: type this and press Enter:

pwd

It prints the folder you are currently in. If it prints a path like /home/yourname, the terminal works. On a server, echo $DISPLAY printing nothing is normal - the default headless browser does not need a screen.

Two words you will see a lot:

  • sudo - "run this as administrator." It will ask for your login password. Nothing appears on screen while you type the password - that is normal. Type it and press Enter.

  • apt - Ubuntu/Debian's built-in installer for system tools. If you use Fedora, replace apt install with dnf install.


Step 2: Install the Basic Tools (unzip and curl)

Why do we need these?

unzip opens your tiktok.zip file. curl downloads the uv installer. Neither is installed by default on a fresh Linux system, so install them now.

sudo apt update
sudo apt install -y unzip curl

Fedora users:

sudo dnf install -y unzip curl

Verify:

unzip -v
curl --version

Each should print version information. If either says command not found, run the install command again and read the error message.


Step 3: Get the Project (extract into a folder named tiktok)

This step fixes the cd tiktok problem

The zip was built with its files at the top level - meaning there is no tiktok folder inside the zip. If you just double-click "Extract Here," the files (main.py, api.py, ...) get dumped loose into the current folder, and later cd tiktok fails with No such file or directory.

The fix: create a folder named tiktok first, then tell unzip to put the files inside it.

# 1. Go to where the uploaded zip lives (usually your home folder)
cd ~

# 2. Create the destination folder
mkdir -p tiktok

# 3. Extract the zip's contents INTO that folder (-d means destination)
unzip -o tiktok.zip -d tiktok

# 4. Enter the folder
cd tiktok

# 5. Confirm the files are here
ls

After step 5, ls must show files such as main.py, pyproject.toml, tiktok_parse.py, and folders browser, models, data.

Already extracted the files loose somewhere? Move them into a proper folder:

mkdir -p ~/tiktok
mv api.py constants.py fetch_sample.py main.py pyproject.toml \
   tiktok_parse.py uv.lock sample.html index.html TUTORIAL.md \
   browser models data ~/tiktok/
cd ~/tiktok

Errors about files that are "not found" are fine if those files are already in place.

If unzip says cannot find or open tiktok.zip: the zip is not in your home folder. Find it with find ~ -name "tiktok.zip", then cd into the folder it printed before running unzip again.

From now on, every command is run while inside ~/tiktok. If you ever get lost, run cd ~/tiktok to return. Confirm with pwd - it should end in /tiktok.


Step 4: Install uv (Package Manager)

This step fixes the uv sync problem (and means you never install Python by hand)

uv is a fast tool that installs the Python libraries this project needs. It also provides Python itself, so you do not install Python - and you should not try to. The python3.12 apt package does not exist on Ubuntu 22.04 / 26.04 or Debian 12 / 13, so trying to install it just fails. uv avoids all of that.

curl -LsSf https://astral.sh/uv/install.sh | sh

This uses curl (installed in Step 2) to download the installer and run it. It puts uv in the folder ~/.local/bin.

Now close and reopen your terminal so it can find the new program. (Or run source $HOME/.local/bin/env.)

Verify:

uv --version

Should show something like uv 0.x.x.

If uv: command not found still appears, add the folder to your PATH, then reload it:

echo 'export PATH="$HOME/.local/bin:$PATH"' >> ~/.bashrc
source ~/.bashrc
uv --version

Step 5: Install a Browser

Brave or Google Chrome

The program controls a real browser through a debugging port. Either browser works.

Tested on Brave. This scraper was built and verified with Brave. It is the recommended browser and the default path (/usr/bin/brave-browser) already points at it.

Brave (recommended) - official repository:

sudo apt update
sudo apt install -y curl gpg

sudo curl -fsSLo /usr/share/keyrings/brave-browser-archive-keyring.gpg \
  https://brave-browser-apt-release.s3.brave.com/brave-browser-archive-keyring.gpg

echo "deb [signed-by=/usr/share/keyrings/brave-browser-archive-keyring.gpg] \
https://brave-browser-apt-release.s3.brave.com/ stable main" \
  | sudo tee /etc/apt/sources.list.d/brave-browser-release.list

sudo apt update
sudo apt install -y brave-browser

More detail is available at brave.com/linux.

Google Chrome instead:

sudo apt install -y wget
wget https://dl.google.com/linux/direct/google-chrome-stable_current_amd64.deb
sudo apt install -y ./google-chrome-stable_current_amd64.deb

Installing this .deb is what creates /usr/bin/google-chrome. If Chrome is installed some other way, that path may not exist.

Find the browser's location. The default in constants.py is /usr/bin/brave-browser. Check the real path with:

which brave-browser
which google-chrome

If it prints a different path, you will edit one line of constants.py later (the file is in the reference section).

MUST-DO - required for Step 8. Headless fetching is impossible without this. Even though the final fetch is headless, the mandatory first run (Step 8A) is headed and needs a virtual screen on a server. Install it now:

sudo apt install -y xvfb
which xvfb-run
xvfb-run --help

which xvfb-run must print a path (usually /usr/bin/xvfb-run). Fedora users: sudo dnf install -y xorg-x11-server-Xvfb. Step 8A will not run without this.


Step 6: Install the Project's Packages

Run uv sync from inside the project

Make sure you are in the project folder first, then sync:

cd ~/tiktok
uv sync

This reads pyproject.toml and installs beautifulsoup4, lxml, httpx, pydantic, psutil and websocket-client, plus developer tools (pytest, ruff, pyright).

If uv sync fails with uv: command not found: go back to Step 4. uv sync cannot work until uv --version works.

If uv sync fails with network/download errors: check your internet connection and run uv sync again. The first run downloads everything.

Linux only - pyright dependency: If uv run pyright fails with libatomic.so.1: cannot open shared object file:

sudo apt install -y libatomic1

Step 7: Run It (offline first)

First run - parse the saved page (no browser)

uv run python main.py

What happens:

  1. A sample.html ships with the project, so it is parsed directly and no browser opens

  2. The parser reads the HTML and prints each video

Runs after that: the browser never opens unless you ask with --fetch. The program reads sample.html and parses it instantly.

Expected Output

Using cached /home/<you>/tiktok/sample.html (533498 bytes)

============================================================
SEARCH: 'new trend'  (12 videos)
============================================================

   1. @eyochico (grimes)
      url      : https://www.tiktok.com/@eyochico/video/7585351064929553677
      video id : 7585351064929553677
      desc     : holy duo #foryou #adrianalima #meganfox #mogging #targetaudience
      likes    : 237800   comments: -   shares: -   plays: -
      posted   : 2025-12-19
      hashtags : #foryou #adrianalima #meganfox #mogging #targetaudience
      cover    : https://p16-common-sign.tiktokcdn.com/tos-useast5-p-0068-tx/...

============================================================
Done.

The number of videos changes from one capture to the next. The sample bundled in this zip holds 12 videos. A live fetch done as 8A headed-init (must fail) → 8B headless retry returns at most 11/12 videos. The SEARCH: line always prints the real number. The first item is the same (@eyochico) across samples of this query.

Want to search something else?

# A different query always needs --fetch — redo 8A then 8B:
xvfb-run -a uv run python main.py --query "dance challenge" --fetch --headed
# expect 0 videos (init), then:
uv run python main.py --query "dance challenge" --fetch
# if 0 videos, run the same headless command once more (max ~11/12)

If you run --query "something" without --fetch, the program stops and tells you there is no downloaded cache for it. That is on purpose - it will not show you the wrong query's data.


Step 8: Live Fetch (get fresh results) — do 8A first, then 8B

Order is mandatory. Do not reverse it. Steps 1-7 were normal setup. Now: 8A headed-under-Xvfb first (it MUST fail), then 8B headless (retry for ~11/12 videos). Running headless before 8A will not work.

Step 8A — Init run (MUST fail with 0 videos, ~90% of the time)

xvfb-run -a uv run python main.py --fetch --headed

This step does not give you videos. Its whole purpose is to invoke TikTok's JavaScript to verify your session and cookies and initialize the page/profile. Expect:

  1. All scrolls show 0 video links rendered

  2. Program stops with Fetch returned 0 videos

  3. Failed page goes to sample.failed.html — your good sample.html is not overwritten

If 8A shows 0 videos, it worked as intended. Move to 8B. This requires xvfb from Step 5 — without xvfb-run this step (and therefore headless) is impossible.

Step 8B — Real fetch (headless, only AFTER 8A has failed)

# Run only after 8A has failed once:
uv run python main.py --fetch

A Brave/Chrome window stays invisible (headless) - that is normal. It loads the page, waits about 12 seconds, scrolls, saves sample.html, then closes.

8B still fails the first time — this is also expected. If the first 8B run returns 0 videos, run the exact same headless command a second time:

uv run python main.py --fetch

The second try generates at most 11/12 videos. That is success. Do not expect 47-83. If it is still empty after two 8B tries, see Troubleshooting.

You will see scroll messages like:

  scroll 1/6: ~12 video links rendered
  scroll 2/6: ~31 video links rendered

In 8A every scroll showing 0 is expected (init failure). In 8B retry the first scroll already shows a non-zero number (about 11/12 headless). If 8B still shows 0 on every scroll twice in a row, you got a login wall/captcha — see Troubleshooting.


Step 9: Share the Browser Profile (optional)

Reuse an existing logged-in browser session

If you also have the Instagram project, you can share its browser profile so both projects use the same cookies/session. Skip this whole step if you do not have the Instagram project - the browser creates its own profile on first run.

# run from inside ~/tiktok
cd ~/tiktok
mkdir -p ../instagram/browser/browser_profile
ln -s ../../instagram/browser/browser_profile browser/browser_profile
ls -l browser/browser_profile        # verify the arrow points to a real folder

If the link already exists (or is broken) you get a FileExistsError. Remove it and create it again:

rm browser/browser_profile
ln -s ../../instagram/browser/browser_profile browser/browser_profile

Troubleshooting

Common Problems and Fixes

cd tiktok - "No such file or directory"

The zip has no tiktok folder inside it. Redo Step 3: mkdir -p tiktok, then unzip -o tiktok.zip -d tiktok, then cd tiktok. Run ls to see where you are.

"Command not found: uv" or uv sync fails

You skipped Step 4. Install uv with the curl command, close and reopen the terminal, and check uv --version. If it still fails, add ~/.local/bin to your PATH (Step 4).

"Command not found: python3.12"

You do not need system Python. uv supplies Python itself. Do not run apt install python3.12 - that package does not exist on Ubuntu 22.04 / 26.04 or Debian 12 / 13.

"Command not found: curl"

Install it: sudo apt install -y curl (Step 2).

"Command not found: unzip"

Install it: sudo apt install -y unzip (Step 2).

Browser doesn't open / "connection refused" / hangs on "Waiting for browser to be ready"

The program now times out instead of hanging forever and prints the browser log. Check constants.py: find your browser path with which brave-browser or which google-chrome and set DEBUG_BROWSER_PATH to match.

Error mentions Missing X server or $DISPLAY

You ran Step 8A without Xvfb. Install it: sudo apt install -y xvfb (Step 5) and always prefix headed with xvfb-run -a: xvfb-run -a uv run python main.py --fetch --headed. Verify with which xvfb-run.

Chrome doesn't work

The scraper was tested on Brave. Install Brave (Step 5) and set DEBUG_BROWSER_PATH = "/usr/bin/brave-browser" in constants.py. Note: /usr/bin/google-chrome only exists if you installed Chrome from the .deb.

0 videos on Step 8A (--headed under xvfb-run)

This MUST happen (~90%). 8A is init-only: it invokes TikTok JavaScript to verify session/cookies. Do not retry 8A — move to Step 8B headless: uv run python main.py --fetch.

0 videos on the first 8B (headless) run

Also expected. Run the exact same headless command a second time: uv run python main.py --fetch. The second try generates at most 11/12 videos.

0 videos even after two 8B runs

Look at sample.failed.html (the saved error page). It may be a login wall or captcha, or your network/profile is blocked. Your sample.html is safe; restore the previous one with cp sample.html.bak sample.html if needed.

Only 11/12 videos (or fewer)

That is success — headless caps at ~11/12 after the 8A init + 8B retry sequence. Do not switch back to --headed for more; if you get fewer than 11, increase SCROLL_STEPS and SCROLL_PAUSE in constants.py, then redo 8A-fail → 8B-retry with --fetch.

Port 9999 already in use

Another browser is still open. Close it and try again, or pass a different debug_port to TikTokClient in your own script (the port is honored, not hard-coded).

FileExistsError on browser_profile

The symlink is broken. Remove it and recreate it (Step 9).

"No module named 'models'"

Make sure you ran unzip -o tiktok.zip -d tiktok fully, so the models/ and browser/ folders exist. Confirm with ls.

pytest says "no tests ran"

Expected - the zip ships no tests/ folder. Not a failure.


Reference: What's Inside the Project

Everything below already ships inside tiktok.zip - you do not need to create, copy, or paste any of it. This section is here only so you know what each file does and can see the code if you are curious. None of it comes from the Instagram project.

pyproject.toml

[project]
name = "tiktok-scraper"
version = "0.1.0"
description = "TikTok search result scraper using Brave CDP browser"
readme = "README.md"
requires-python = ">=3.12"
dependencies = [
    "beautifulsoup4>=4.15.0",
    "httpx>=0.28.1",
    "lxml>=6.1.1",
    "psutil>=7.0.0",
    "pydantic>=2.13.4",
    "websocket-client>=1.9.0",
]

[dependency-groups]
dev = [
    "pyright>=1.1.410",
    "pytest>=9.0.3",
    "ruff>=0.15.16",
]

[tool.ruff]
target-version = "py312"
line-length = 88

[tool.ruff.lint]
fixable = ["ALL"]
select = ["I", "B", "E"]

[tool.pyright]
typeCheckingMode = "strict"
venvPath = "."
venv = ".venv"
pythonVersion = "3.12"

[tool.pytest.ini_options]
testpaths = ["tests"]
pythonpath = ["."]

constants.py

from pathlib import Path

# Paths
HOME = Path.home()
BASE_DIR = Path(__file__).parent

# Browser
DEBUG_BROWSER_PATH = "/usr/bin/brave-browser"
# Symlinked to ../../instagram/browser/browser_profile so the TikTok runs reuse
# the same browser session / cookies as the Instagram project.
USER_PROFILE_DIR = BASE_DIR / "browser" / "browser_profile"

# Browser Configs
WIN_W = 720
WIN_H = 760

# TikTok
TIKTOK_BASE = "https://www.tiktok.com"
DEFAULT_QUERY = "new trend"
SAMPLE_HTML_PATH = BASE_DIR / "sample.html"

# Kept so a failed live fetch never destroys a good sample.html.
SAMPLE_HTML_BACKUP_PATH = BASE_DIR / "sample.html.bak"
# Remembers which query produced sample.html, so a mismatched --query is caught.
SAMPLE_META_PATH = BASE_DIR / "sample.meta.json"

# Browser stdout/stderr are captured here instead of being thrown away.
BROWSER_LOG_PATH = BASE_DIR / "browser" / "browser.log"
# Seconds to wait for the browser debug port before giving up.
BROWSER_READY_TIMEOUT = 30

# How long to let the client-rendered search page settle before grabbing HTML.
# Note: headless mode may render fewer results than headed mode. Use --headed
# (under xvfb-run on a server) when you need the full result set.
RENDER_TIMEOUT = 12

# Scroll settings used to trigger lazy-loaded search results.
SCROLL_STEPS = 6
SCROLL_DELTA_Y = 4000
SCROLL_PAUSE = 2.0

If your browser is not at /usr/bin/brave-browser, edit the DEBUG_BROWSER_PATH line. Find the path with which brave-browser or which google-chrome.

models/search.py - the data shape

from pydantic import BaseModel, Field


class TikTokVideoItem(BaseModel):
    """One video item parsed from a TikTok search result page.

    All fields are optional / defaulted so a card can still be captured even
    when the rendered DOM omits a specific value.

    Note: TikTok renders the engagement number inside the search card with a
    heart (like) icon but labels the element ``data-e2e="video-views"``. It is
    mapped to :attr:`likes`. Comments / shares / plays are not present in the
    search-result DOM, so they stay ``None`` unless a future layout adds them.
    """

    video_id: str | None = None
    video_url: str | None = None
    description: str | None = None
    hashtags: list[str] = Field(default_factory=list)

    author_username: str | None = None
    author_nickname: str | None = None
    author_url: str | None = None
    author_avatar_url: str | None = None
    author_verified: bool | None = None

    likes: int | None = None
    comments: int | None = None
    shares: int | None = None
    plays: int | None = None

    cover_url: str | None = None
    duration: int | None = None
    create_time: int | None = None
    create_time_text: str | None = None

    music_title: str | None = None
    top_liked: bool = False

    likes_raw: str | None = None
    comments_raw: str | None = None
    shares_raw: str | None = None
    plays_raw: str | None = None


class TikTokSearchData(BaseModel):
    query: str | None = None
    result_count: int = 0
    videos: list[TikTokVideoItem] = Field(default_factory=list[TikTokVideoItem])

tiktok_parse.py - how the page is read

This finds the result list and reads each card. Its core selectors:

SelectorWhat it is
[data-e2e="search_top-item-list"]the results container
[data-e2e="search_top-item"]one video card (link + likes)
[data-e2e="search-card-desc"]caption + author block
[data-e2e="search-card-video-caption"]caption text + hashtags
[data-e2e="search-card-user-link"]author link (/@username)
[data-e2e="search-card-user-unique-id"]author nickname
[data-e2e="video-views"]the likes number (next to a heart icon)

Counts are abbreviated. TikTok writes 39.4K or 1.2M. The parser converts them to 39400 and 1200000 so you can sort and add them up.

Full source:

"""
Parse a rendered TikTok search result page into structured data.

Extraction strategy (matching the rendered DOM served today):
- Each result lives in a direct child of the ``[data-e2e="search_top-item-list"]``
  container.
- The video link/cover/likes come from the ``[data-e2e="search_top-item"]`` card.
- The caption, hashtags and author come from the sibling
  ``[data-e2e="search-card-desc"]`` block.

TikTok search pages are client-side rendered: the server HTML is only a shell.
This parser therefore expects the *browser-rendered* DOM saved to sample.html by
``fetch_sample.py``.

Run directly to validate against the local sample.html:
  uv run python tiktok_parse.py
"""

import json
import re

from bs4 import BeautifulSoup
from bs4.element import Tag

from models.search import TikTokSearchData, TikTokVideoItem

TIKTOK_BASE = "https://www.tiktok.com"

LIST_SELECTOR = '[data-e2e="search_top-item-list"]'
CARD_SELECTOR = '[data-e2e="search_top-item"]'
CAPTION_SELECTOR = '[data-e2e="search-card-video-caption"]'
USER_LINK_SELECTOR = '[data-e2e="search-card-user-link"]'
NICKNAME_SELECTOR = '[data-e2e="search-card-user-unique-id"]'
LIKE_SELECTOR = '[data-e2e="video-views"]'
HASHTAG_SELECTOR = '[data-e2e="search-common-link"]'

_VIDEO_ID_RE = re.compile(r"/video/(\d+)")
_COUNT_RE = re.compile(r"([\d.]+)\s*([KMB]?)", re.IGNORECASE)
_COUNT_MULT = {"": 1, "K": 1_000, "M": 1_000_000, "B": 1_000_000_000}


def parse_count(text: str | None) -> int | None:
    """Convert TikTok's abbreviated counts ("1.2M", "39.4K", "1760") to int."""
    if not text:
        return None
    cleaned = text.replace(",", "").replace(" ", "").strip()
    match = _COUNT_RE.search(cleaned)
    if not match:
        return None
    try:
        value = float(match.group(1))
    except ValueError:
        return None
    multiplier = _COUNT_MULT.get(match.group(2).upper(), 1)
    return int(value * multiplier)


def _select_one(node: Tag, selector: str) -> Tag | None:
    return node.select_one(selector)


def _select(node: Tag, selector: str) -> list[Tag]:
    return node.select(selector)


def _attr(node: Tag, name: str) -> str | None:
    value = node.get(name)
    return value if isinstance(value, str) else None


def _text(node: Tag | None) -> str | None:
    if node is None:
        return None
    value = node.get_text(" ", strip=True)
    return value or None


def _parse_hashtags(caption: Tag | None) -> list[str]:
    if caption is None:
        return []
    tags: list[str] = []
    for link in _select(caption, HASHTAG_SELECTOR):
        href = _attr(link, "href") or ""
        if href.startswith("/tag/"):
            tag = link.get_text(strip=True)
            if tag and tag not in tags:
                tags.append(tag)
    return tags


def _parse_time_text(user_link: Tag | None) -> str | None:
    """The relative/absolute post date sits next to the author link."""
    if user_link is None:
        return None
    row = user_link.parent
    if not isinstance(row, Tag):
        return None
    for sibling in row.find_all(recursive=False):
        if sibling is user_link:
            continue
        text = sibling.get_text(strip=True)
        if text:
            return text
    return None


def _parse_video_link(card: Tag) -> Tag | None:
    for anchor in card.find_all("a"):
        href = _attr(anchor, "href") or ""
        if _VIDEO_ID_RE.search(href):
            return anchor
    return None


def parse_video_item(container: Tag) -> TikTokVideoItem | None:
    """Parse one search result container into a TikTokVideoItem."""
    card = _select_one(container, CARD_SELECTOR)
    if card is None:
        return None

    link = _parse_video_link(card)
    if link is None:
        return None

    video_url = _attr(link, "href")
    video_match = _VIDEO_ID_RE.search(video_url or "")
    video_id = video_match.group(1) if video_match else None

    cover = _select_one(link, "img")
    cover_url = _attr(cover, "src") if cover is not None else None

    likes_raw = _text(_select_one(card, LIKE_SELECTOR))

    caption = _select_one(container, CAPTION_SELECTOR)
    description = _text(caption)
    hashtags = _parse_hashtags(caption)

    user_link = _select_one(container, USER_LINK_SELECTOR)
    author_username = None
    author_url = None
    author_avatar_url = None
    if user_link is not None:
        href = _attr(user_link, "href") or ""
        if href.startswith("/@"):
            author_username = href[2:]
            author_url = f"{TIKTOK_BASE}{href}"
        avatar = _select_one(user_link, "img")
        author_avatar_url = _attr(avatar, "src") if avatar is not None else None

    author_nickname = _text(_select_one(container, NICKNAME_SELECTOR))
    create_time_text = _parse_time_text(user_link)

    top_liked = "Top liked" in card.get_text()

    return TikTokVideoItem(
        video_id=video_id,
        video_url=video_url,
        description=description,
        hashtags=hashtags,
        author_username=author_username,
        author_nickname=author_nickname,
        author_url=author_url,
        author_avatar_url=author_avatar_url,
        likes=parse_count(likes_raw),
        likes_raw=likes_raw,
        cover_url=cover_url,
        create_time_text=create_time_text,
        top_liked=top_liked,
    )


def _extract_query(soup: BeautifulSoup) -> str | None:
    box = soup.select_one('[data-e2e="search-user-input"]')
    if box is not None:
        value = box.get("value")
        if isinstance(value, str) and value:
            return value
    return None


def parse_search_page(html: str, query: str | None = None) -> TikTokSearchData:
    """Parse rendered TikTok search HTML and return structured data."""
    soup = BeautifulSoup(html, "lxml")

    container = soup.select_one(LIST_SELECTOR)
    items: list[TikTokVideoItem] = []
    seen_ids: set[str] = set()

    if container is not None:
        for child in container.find_all(recursive=False):
            item = parse_video_item(child)
            if item is None or item.video_id is None:
                continue
            if item.video_id in seen_ids:
                continue
            seen_ids.add(item.video_id)
            items.append(item)

    return TikTokSearchData(
        query=query or _extract_query(soup),
        result_count=len(items),
        videos=items,
    )


if __name__ == "__main__":
    from constants import SAMPLE_HTML_PATH

    if not SAMPLE_HTML_PATH.exists():
        print("sample.html not found. Run fetch_sample.py first.")
        raise SystemExit(1)

    data = parse_search_page(SAMPLE_HTML_PATH.read_text(encoding="utf-8"))

    print("=== Search ===")
    print(f"  query        : {data.query}")
    print(f"  results      : {data.result_count}")
    print()
    for i, video in enumerate(data.videos, 1):
        print(f"  {i:>2}. @{video.author_username}  ({video.author_nickname})")
        print(f"      url      : {video.video_url}")
        print(f"      desc     : {(video.description or '')[:80]}")
        print(f"      likes    : {video.likes}  (raw {video.likes_raw!r})")
        print(f"      posted   : {video.create_time_text}")
        print(f"      hashtags : {video.hashtags}")
        print(f"      cover    : {(video.cover_url or '')[:60]}...")
    print()
    print(json.dumps(data.model_dump(mode="json"), ensure_ascii=False)[:500] + " ...")

api.py and main.py - fetching and the command line

api.py drives the browser (or reads sample.html) and calls the parser. main.py is the command-line tool with --query, --fetch, and --headed. Full source:

"""
TikTok search scraper client.

TikTok search pages are client-side rendered, so live fetches go through the
Brave CDP browser and capture the rendered DOM after scrolling to trigger
lazy-loaded results.

If sample.html already exists, the browser is never spawned - all parse work
runs against the local file.

The browser runs **headless** by default. On a fresh browser profile the very
first live fetch returns 0 videos (TikTok serves an error page instead of a
login wall or captcha). In that case the fetch is stopped with
:class:`FirstRunWarmupError`, which tells the caller to rerun the same command
once. The first run only warms cookies/caches.
"""

import json
import time
from pathlib import Path
from urllib.parse import quote

from browser.actions import (
    browser_open_url,
    dom_element_html,
    dom_js_execute,
    dom_scroll_at,
)
from browser.browser import get_browser_instance, kill_browser
from constants import (
    RENDER_TIMEOUT,
    SAMPLE_HTML_BACKUP_PATH,
    SAMPLE_META_PATH,
    SCROLL_DELTA_Y,
    SCROLL_PAUSE,
    SCROLL_STEPS,
    TIKTOK_BASE,
    WIN_H,
    WIN_W,
)
from models.search import TikTokSearchData
from tiktok_parse import parse_search_page

VIDEO_LINK_COUNT_JS = 'document.querySelectorAll(\'a[href*="/video/"]\').length'

FIRST_RUN_MESSAGE = (
    "Fetch returned 0 videos (0/0). The first fetch on a fresh browser profile "
    "only warms cookies/caches and is expected to fail on this same system. "
    "Rerun the exact same command once - it should then return results, e.g.:\n"
    "    uv run python main.py --fetch\n"
    "Do not change code or config. A second run normally returns videos; if it "
    "still returns 0, the profile/network is the problem, not the setup."
)


class FirstRunWarmupError(RuntimeError):
    """Raised when a live fetch returns zero videos on a cold profile."""


class TikTokClient:
    def __init__(self, debug_port: int = 9999, headless: bool = True) -> None:
        self.debug_port = debug_port
        self.headless = headless
        self._browser = None

    def _ensure_browser(self) -> None:
        if self._browser is None:
            mode = "headless" if self.headless else "headed"
            print(f"Spawning Brave browser ({mode}) ...")
            self._browser = get_browser_instance(
                self.debug_port, headless=self.headless
            )

    def _scroll_to_load(self) -> None:
        assert self._browser is not None
        cx, cy = WIN_W // 2, WIN_H // 2
        for i in range(SCROLL_STEPS):
            dom_scroll_at(self._browser, cx, cy, delta_y=SCROLL_DELTA_Y)
            time.sleep(SCROLL_PAUSE)
            count = dom_js_execute(self._browser, VIDEO_LINK_COUNT_JS).get("value", 0)
            print(f"  scroll {i + 1}/{SCROLL_STEPS}: ~{count} video links rendered")

    @staticmethod
    def search_url(query: str) -> str:
        return f"{TIKTOK_BASE}/search?q={quote(query)}"

    def fetch_search_html(self, query: str) -> str:
        """Navigate straight to the search page and return the rendered DOM.

        There is deliberately no visit to tiktok.com first and no in-process
        retry: a cold-profile failure is reported to the caller so the operator
        reruns once, which is all that is needed.
        """
        url = self.search_url(query)
        self._ensure_browser()
        assert self._browser is not None
        print(f"Navigating to {url} ...")
        browser_open_url(self._browser, url)
        print(f"Waiting {RENDER_TIMEOUT}s for the client-rendered page ...")
        time.sleep(RENDER_TIMEOUT)
        self._scroll_to_load()
        time.sleep(SCROLL_PAUSE)
        return dom_element_html(self._browser, "body", outer_html=True)

    @staticmethod
    def _save_sample(html_path: Path, query: str, html: str) -> None:
        if html_path.exists():
            html_path.replace(SAMPLE_HTML_BACKUP_PATH)
            print(
                f"Backed up previous {html_path.name} to "
                f"{SAMPLE_HTML_BACKUP_PATH.name}"
            )
        html_path.write_text(html, encoding="utf-8")
        SAMPLE_META_PATH.write_text(json.dumps({"query": query}), encoding="utf-8")
        print(f"Saved {len(html)} bytes to {html_path}")

    def scrape_search(
        self,
        query: str,
        *,
        html_path: Path | None = None,
        force: bool = False,
    ) -> TikTokSearchData:
        """
        Return parsed search data for query.

        If html_path exists and force is False, parse the local file (no
        browser). Otherwise fetch via the browser, save to html_path, then parse.

        Raises FirstRunWarmupError when a live fetch returns 0 videos, so the
        existing sample.html is left untouched.
        """
        if html_path and html_path.exists() and not force:
            html = html_path.read_text(encoding="utf-8")
            return parse_search_page(html, query=query)

        html = self.fetch_search_html(query)
        data = parse_search_page(html, query=query)

        if data.result_count == 0:
            if html_path:
                failed_path = html_path.with_suffix(".failed.html")
                failed_path.write_text(html, encoding="utf-8")
                print(f"Saved the failed fetch to {failed_path} for inspection")
            raise FirstRunWarmupError(FIRST_RUN_MESSAGE)

        if html_path:
            self._save_sample(html_path, query, html)
        return data

    def close(self) -> None:
        if self._browser:
            kill_browser(self._browser)
            self._browser = None

    def __enter__(self) -> "TikTokClient":
        return self

    def __exit__(
        self,
        exc_type: type[BaseException] | None,
        exc: BaseException | None,
        tb: object | None,
    ) -> None:
        self.close()
"""
TikTok search result scraper.

Default: parse sample.html (local, no browser). Live fetches run **headless**
by default; pass --headed to open a visible window (needs a display or Xvfb).

Usage:
  uv run python main.py                              # offline, parse sample.html
  uv run python main.py --fetch                      # live fetch (headless) + parse
  uv run python main.py --fetch --headed             # live fetch in a window
  uv run python main.py --query "new trend" --fetch  # fetch a different query

The first live fetch on a fresh browser profile returns 0 videos on purpose
(it only warms cookies/caches). Rerun the same command once and it works.
"""

import argparse
import json
import sys

from api import FirstRunWarmupError, TikTokClient
from constants import DEFAULT_QUERY, SAMPLE_HTML_PATH, SAMPLE_META_PATH


def print_separator(char: str = "=", length: int = 60) -> None:
    print(char * length)


def _count(value: int | None) -> str:
    return str(value) if value is not None else "-"


def _cached_query() -> str:
    """Query that produced the bundled/cached sample.html (default fallback)."""
    if SAMPLE_META_PATH.exists():
        try:
            meta = json.loads(SAMPLE_META_PATH.read_text(encoding="utf-8"))
            query = meta.get("query")
            if isinstance(query, str) and query:
                return query
        except (OSError, ValueError):
            pass
    return DEFAULT_QUERY


def main() -> None:
    parser = argparse.ArgumentParser(description="TikTok search scraper")
    parser.add_argument(
        "--query",
        default=DEFAULT_QUERY,
        help=f"search query (default: {DEFAULT_QUERY!r}); needs --fetch if it "
        "differs from the cached sample",
    )
    parser.add_argument(
        "--fetch",
        action="store_true",
        help="Force a live browser fetch even if sample.html exists",
    )
    parser.add_argument(
        "--headed",
        action="store_true",
        help="Open a visible browser window instead of headless (needs a "
        "display; on a server use `xvfb-run -a ... --headed`)",
    )
    args = parser.parse_args()

    html_path = SAMPLE_HTML_PATH
    cached_query = _cached_query()

    if not args.fetch and args.query != cached_query:
        print(
            f"No downloaded cache for query {args.query!r} "
            f"(the cached sample is for {cached_query!r}).\n"
            "Add --fetch to download it:\n"
            f"    uv run python main.py --query {args.query!r} --fetch\n"
            "Expect the first run to return 0 videos - if so, rerun the exact "
            "same command once. This is normal on a fresh browser profile.",
            file=sys.stderr,
        )
        raise SystemExit(2)

    if html_path.exists() and not args.fetch:
        print(f"Using cached {html_path} ({html_path.stat().st_size} bytes)")
    else:
        print("Fetching from browser ...")

    try:
        with TikTokClient(headless=not args.headed) as client:
            data = client.scrape_search(
                args.query, html_path=html_path, force=args.fetch
            )
    except FirstRunWarmupError as exc:
        print(f"\n{exc}", file=sys.stderr)
        raise SystemExit(1) from exc
    except RuntimeError as exc:
        print(f"\nBrowser startup failed:\n{exc}", file=sys.stderr)
        raise SystemExit(1) from exc

    print()
    print_separator()
    print(f"SEARCH: {data.query!r}  ({data.result_count} videos)")
    print_separator()

    for i, video in enumerate(data.videos, 1):
        print(f"\n  {i:>2}. @{video.author_username} ({video.author_nickname})")
        print(f"      url      : {video.video_url}")
        print(f"      video id : {video.video_id}")
        print(f"      desc     : {(video.description or '-')[:90]}")
        print(
            f"      likes    : {_count(video.likes)}"
            f"   comments: {_count(video.comments)}"
            f"   shares: {_count(video.shares)}"
            f"   plays: {_count(video.plays)}"
        )
        print(f"      posted   : {video.create_time_text or '-'}")
        if video.hashtags:
            print(f"      hashtags : {' '.join(video.hashtags)}")
        if video.cover_url:
            print(f"      cover    : {video.cover_url[:70]}...")

    print()
    print_separator()
    print("Done.")


if __name__ == "__main__":
    main()

browser/ - the browser control layer

These files (browser.py, actions.py, utils.py) open Brave/Chrome with a debugging port, connect over CDP, navigate, scroll, and grab the finished HTML. They ship ready to use.


Final File Checklist

Your ~/tiktok folder should contain exactly this:


tiktok/
├── index.html
├── TUTORIAL.md
├── pyproject.toml
├── uv.lock
├── constants.py
├── api.py
├── tiktok_parse.py
├── fetch_sample.py
├── main.py
├── sample.html
├── browser/
│ ├── init.py
│ ├── browser.py
│ ├── actions.py
│ └── utils.py
├── models/
│ ├── init.py
│ ├── browser.py
│ └── search.py
└── data/
└── .gitkeep

Some extra files appear only after you run things - they are safe to ignore or delete:


sample.html.bak # previous good sample, kept when a new fetch succeeds
sample.failed.html # the error page from a failed (0-video) fetch
sample.meta.json # which query produced sample.html
browser/browser.log # browser messages, useful when something goes wrong

You're done! If every file is in place and you ran uv sync, the scraper works. Run it anytime with:

cd ~/tiktok
uv run python main.py

Questions? Double-check file names and paths - most issues are a missed step above or a typo.

Get Started by signing up for a Proxy Product