About this Dataset
This ChatGPT AI dataset is a free CSV sample that shows how to scrape ChatGPT: a prompt is sent to a Scrape.do endpoint for the ChatGPT plugin, and the returned JSON is parsed into useful, structured fields. It ships as two files — one for the answer plus model, message, and latency metadata, and one for the code blocks inside the answer, split row by row. This free ai dataset download requires no coding and gives you a realistic preview of the ChatGPT data you can collect at scale before committing to a larger project.
The primary file in this ChatGPT dataset includes 42 fields: prompt, answer_text, answer_markdown, answer_html, char_count, word_count, line_count, heading_count, and more. It captures the answer as text, markdown, and HTML, together with content metrics, model and message identifiers (model_slug, message_id, conversation_id), request and latency signals (upstream_latency, upstream_latency_sec), and timestamps. The companion code file breaks every code block out into its own row with 6 fields: prompt, message_id, block_index, language, line_count, code. Both files are normalized into a consistent schema and delivered in CSV format, ready to import into spreadsheets, databases, or analytics and evaluation pipelines.
This sample is especially useful for AI and LLM teams, prompt engineers, researchers, and evaluation platforms that need reliable ai data extraction without building and maintaining their own scrapers. Whether you are benchmarking model answers, mining code snippets, or analyzing latency and response structure, structured ChatGPT data speeds up every step of the workflow.
DataHarbor collected this sample using the same managed infrastructure that powers our custom data extraction across 700M+ domains — sending prompts through a Scrape.do endpoint, handling proxy rotation, rendering, and request management, then parsing the raw JSON response into clean, validated, and deduplicated CSV or Excel output. Only publicly available response information is collected.
Need more than a sample? DataHarbor delivers full-scale ChatGPT scraping — thousands of prompts and responses with the exact fields you need — refreshed daily, weekly, or monthly in CSV, Excel, JSON, or straight to your database. Contact us to request a custom ai dataset built to your exact requirements.
Data Fields
| # | Field Name |
|---|---|
| 1 | prompt |
| 2 | answer_text |
| 3 | answer_markdown |
| 4 | answer_html |
| 5 | char_count |
| 6 | word_count |
| 7 | line_count |
| 8 | heading_count |
| 9 | code_block_count |
| 10 | code_languages |
| 11 | message_id |
| 12 | conversation_id |
| 13 | parent_message_id |
| 14 | author_role |
| 15 | model_slug |
| 16 | resolved_model_slug |
| 17 | status |
| 18 | is_complete |
| 19 | end_turn |
| 20 | message_type |
| 21 | content_type |
| 22 | recipient |
| 23 | channel |
| 24 | finish_type |
| 25 | stop_tokens |
| 26 | request_id |
| 27 | turn_exchange_id |
| 28 | create_time_unix |
| 29 | create_time_utc |
| 30 | update_time_unix |
| 31 | citation_count |
| 32 | source_count |
| 33 | search_query_count |
| 34 | shopping_count |
| 35 | ads_count |
| 36 | stream_bytes |
| 37 | upstream_latency |
| 38 | upstream_latency_sec |
| 39 | geo_code |
| 40 | error |
| 41 | error_code |
| 42 | fetched_at_utc |
Note: The fields included in these sample datasets are selected for demonstration purposes only. For your actual project, we can extract any available data field from the target website — fully customized to match your specific requirements.
Download
how-to-scrape-chatgpt.csv
CSV · 1 records · 42 fields
Code Blocks — chatgpt_responses-code.csv
CSV · 2 records · 6 fields
prompt, message_id, block_index, language, line_count, code
Need More Data?
This is just a sample. DataHarbor delivers full-scale, custom ChatGPT datasets — more records, more fields, more regions — on the schedule you need, in the format you want.