{"id":"dataset-101","title":"Train, validation and test sets","messages":[{"role":"user","content":"I have a small dataset of support conversations. How should I split it for evaluation?"},{"role":"assistant","content":"## Give each split a clear job\n\n- **Train:** examples the model learns from.\n- **Validation:** examples you use to compare changes.\n- **Test:** a final check after you have made your choices.\n\nKeep messages from the **same conversation in one split**. Otherwise, part of a conversation can leak into the evaluation set."},{"role":"user","content":"So I should split by conversation ID, rather than by message?"},{"role":"assistant","content":"Exactly. Assign each conversation ID to one split, then keep all of its messages together.\n\n```python\ntrain = [chat for chat in chats if chat[\"id\"] in train_ids]\ntest = [chat for chat in chats if chat[\"id\"] in test_ids]\n```\n\nAlso check for near-duplicate conversations before evaluating."}]}
{"id":"dataset-102","title":"Reading a saved chat export","messages":[{"role":"user","content":"What should I check before using a downloaded conversation dataset?"},{"role":"assistant","content":"## Read a few complete conversations\n\nCheck whether **roles, turn order and formatting** match what you expect. A row may contain one message or an entire conversation.\n\nLook at short and long examples, code blocks, tool results and incomplete turns. Then review the dataset card for its source, license and filtering."}]}
{"id":"dataset-103","title":"Counting tokens in a JSONL file","messages":[{"role":"system","content":"You are a concise assistant for data engineers."},{"role":"user","content":"How can I estimate the number of tokens in a JSONL chat dataset?"},{"role":"assistant","content":"Sum a tokenizer’s count over every message’s `content`:\n\n```python\nimport json, tiktoken\nenc = tiktoken.get_encoding(\"o200k_base\")\ntotal = 0\nwith open(\"chats.jsonl\") as f:\n    for line in f:\n        for m in json.loads(line)[\"messages\"]:\n            total += len(enc.encode(m[\"content\"] or \"\"))\nprint(total)\n```\n\nCounts vary between tokenizers, so treat the result as an estimate."},{"role":"user","content":"Roughly how many tokens is 1 MB of English text?"},{"role":"assistant","content":"About **250,000 tokens**, at roughly four characters per token. Code and non-English text usually produce more."}]}
{"id":"dataset-104","title":"Removing personal data","messages":[{"role":"user","content":"Our support logs contain email addresses and order numbers. What should I remove before sharing a sample?"},{"role":"assistant","content":"Replace anything that identifies a person or an account:\n\n1. Names, email addresses and phone numbers\n2. Order, ticket and account IDs\n3. Street addresses and payment details\n4. Free-text fields where people describe themselves\n\nUse consistent placeholders like `<EMAIL_1>` so a conversation still reads naturally, and review a random sample by hand afterwards."}]}
