AI

How to Scrape ChatGPT — Responses & Code Dataset

A free sample dataset showing how to scrape ChatGPT: a prompt sent to a Scrape.do endpoint is parsed into structured fields. One file holds the answer plus model, message and latency metadata (42 fields); a second file lists each code block in the answer as its own row.

Source: ChatGPT1 recordsFormat: CSVAugust 12, 2026

About this Dataset

This ChatGPT AI dataset is a free CSV sample that shows how to scrape ChatGPT: a prompt is sent to a Scrape.do endpoint for the ChatGPT plugin, and the returned JSON is parsed into useful, structured fields. It ships as two files — one for the answer plus model, message, and latency metadata, and one for the code blocks inside the answer, split row by row. This free ai dataset download requires no coding and gives you a realistic preview of the ChatGPT data you can collect at scale before committing to a larger project.

The primary file in this ChatGPT dataset includes 42 fields: prompt, answer_text, answer_markdown, answer_html, char_count, word_count, line_count, heading_count, and more. It captures the answer as text, markdown, and HTML, together with content metrics, model and message identifiers (model_slug, message_id, conversation_id), request and latency signals (upstream_latency, upstream_latency_sec), and timestamps. The companion code file breaks every code block out into its own row with 6 fields: prompt, message_id, block_index, language, line_count, code. Both files are normalized into a consistent schema and delivered in CSV format, ready to import into spreadsheets, databases, or analytics and evaluation pipelines.

This sample is especially useful for AI and LLM teams, prompt engineers, researchers, and evaluation platforms that need reliable ai data extraction without building and maintaining their own scrapers. Whether you are benchmarking model answers, mining code snippets, or analyzing latency and response structure, structured ChatGPT data speeds up every step of the workflow.

DataHarbor collected this sample using the same managed infrastructure that powers our custom data extraction across 700M+ domains — sending prompts through a Scrape.do endpoint, handling proxy rotation, rendering, and request management, then parsing the raw JSON response into clean, validated, and deduplicated CSV or Excel output. Only publicly available response information is collected.

Need more than a sample? DataHarbor delivers full-scale ChatGPT scraping — thousands of prompts and responses with the exact fields you need — refreshed daily, weekly, or monthly in CSV, Excel, JSON, or straight to your database. Contact us to request a custom ai dataset built to your exact requirements.

Data Fields

#Field Name
1prompt
2answer_text
3answer_markdown
4answer_html
5char_count
6word_count
7line_count
8heading_count
9code_block_count
10code_languages
11message_id
12conversation_id
13parent_message_id
14author_role
15model_slug
16resolved_model_slug
17status
18is_complete
19end_turn
20message_type
21content_type
22recipient
23channel
24finish_type
25stop_tokens
26request_id
27turn_exchange_id
28create_time_unix
29create_time_utc
30update_time_unix
31citation_count
32source_count
33search_query_count
34shopping_count
35ads_count
36stream_bytes
37upstream_latency
38upstream_latency_sec
39geo_code
40error
41error_code
42fetched_at_utc

Note: The fields included in these sample datasets are selected for demonstration purposes only. For your actual project, we can extract any available data field from the target website — fully customized to match your specific requirements.

Download

how-to-scrape-chatgpt.csv

CSV · 1 records · 42 fields

Download Sample CSV

Code Blocks — chatgpt_responses-code.csv

CSV · 2 records · 6 fields

prompt, message_id, block_index, language, line_count, code

Download CSV

Need More Data?

This is just a sample. DataHarbor delivers full-scale, custom ChatGPT datasets — more records, more fields, more regions — on the schedule you need, in the format you want.

Related Datasets

Amazon
E-commerce

Amazon Coffee Machine Prices Dataset (ZIP 10001)

Featured

A sample price list of coffee machines available on Amazon for ZIP code 10001 (New York), including title, price, list price, rating, review count and delivery options.

48 RecordsJul 25, 2026
eBay
E-commerce

eBay Guitars for Sale Dataset (California)

Featured

A sample dataset of guitars listed for sale in California, scraped from eBay (ebay.com), including listing title, price, condition, brand, location, shipping and detailed product specifications.

60 RecordsJul 25, 2026
Google Search
Search Data

Google Search Results Dataset — Togg (US)

Featured

A sample dataset of Google Search (SERP) results for the query "togg" — the Turkish electric car brand — collected from a US location, including organic result position, title, link, snippet and source.

6 RecordsJul 25, 2026