Video Transcript Knowledge Base: Timestamps, Search and QA

Apify Actors linked under my tugelbay account are my products. I earn from their usage and may also receive a referral commission through Apify links in this article.
How Do You Build a Video Transcript Knowledge Base?
Keep each transcript tied to its source video, track and time range, then test whether a reader can reach the evidence behind an answer. Searchable text is useful, but a paragraph without reliable provenance may send someone to the wrong recording, language or moment.
This tutorial develops a small knowledge-base record from three synthetic caption segments. It includes a complete timestamp-normalization script and a worked question with an expected answer. The transcript is authored for this guide; no real video was downloaded, transcribed or evaluated to create the example. The workflow does not promise caption availability, a product entitlement or a particular transcription accuracy.
The useful deliverable is a source-linked answer with an inspectable segment range, plus a queue for unavailable or uncertain material. You can establish that contract before choosing a search database or committing a whole video library to a paid tool.
Define the Library and the Questions It Should Answer
Start with an authorized collection and a bounded retrieval task. A public educational library, internal training archive and customer-call repository have different permissions, retention needs and risks. Do not merge them into one index merely because each has text.
Write a collection manifest listing the source, permission basis, intended audience, required language and expected output. A URL is an identifier, not proof that collection or redistribution is permitted. Public visibility, unlisted availability and authorized access are different conditions.
Choose representative questions before ingestion. For a training library, ask about concrete steps, prerequisites and exceptions. Include at least one question the library cannot answer. If every evaluation question has a convenient answer, you will not discover whether the system knows when its evidence is insufficient. The human-led AI content workflow provides a related source-note and review pattern when a retrieved passage later becomes published content.
Define whether the output should locate a passage, summarize a lesson or produce an exact quotation. These require different acceptance checks. A search result may point to an approximate segment; an attributed quotation needs the wording checked against the actual recording. Captions can contain mistakes even when their time codes look valid.
For this tutorial, the fictional library has one authored export-training lesson. The question is: “Can I enable exports before the workspace review?” The required evidence includes a negation and the later instruction that resolves the condition. A result that retrieves only “enable exports” is not enough.
Keep Video Identity and Track Identity Separate
A video can have multiple text representations, and they are not interchangeable. Store the selected language, caption origin and any translation status rather than keeping a single unexplained transcript field.
Use explicit unknown or null when the supplied output does not establish a field. Do not infer human authorship from polished text, automatic generation from a typo or speaker identity from sentence style. If a provider gives only plain text, record that limitation instead of inventing segment timing.
| Field | Reason to preserve it | Example for this tutorial |
|---|---|---|
video_id | Stable source identity | fixture-export-training |
source_kind | Separate authored material from captured evidence | authored_fixture |
track_id | Identify the particular text track | fixture-en-v1 |
language | Avoid unnoticed language substitution | en, authored |
caption_origin | Manual, automatic or unknown when established | authored_fixture |
translation_of | Distinguish translation from original wording | null |
video_duration_seconds | Support bounds checks when known | null, no actual recording |
retrieved_at | Timestamp for actual acquisition | null, no acquisition |
Keep the original output with the normalized representation when retention permissions allow it. A corrected track is a new version with a review history. It should not silently overwrite the evidence used by an earlier quotation or answer.
My YouTube Transcript Scraper on Apify is one candidate collection route to inspect. It is my product; I earn from usage and may receive a referral commission through this link. This tutorial does not establish its current inputs, available tracks, output fields or accuracy. Check the current contract and a permitted sample before use. An existing authorized caption export may already meet your needs.
A Completed Synthetic Transcript Example
Preserve the qualification and the action that follows it. The three authored segments below describe an imaginary training lesson, not a real service’s export rules.
| Segment | Original timing contract | Text |
|---|---|---|
seg-01 | Start 12,500 milliseconds; duration 4,000 milliseconds | Do not enable exports before the workspace review. |
seg-02 | Start 16.5 seconds; end 22 seconds | After approval, the administrator can enable exports. |
seg-03 | Start 22 seconds; duration 3 seconds | Keep the review record with the workspace configuration. |
After normalization, the ranges are 12.5–16.5 seconds, 16.5–22 seconds and 22–25 seconds. The first and second segments should remain available together for the intended question. The third adds an operational follow-up but does not cancel the prerequisite.
The reference answer is: “In this fictional lesson, exports should remain disabled until the workspace review. After approval, an administrator can enable them. Evidence: fixture-en-v1, segments 01–02, 12.5–22 seconds.” This is a completed authored answer, not an observed output from a retrieval model.
The rejected answer is: “Enable exports before the review.” It uses words present in the source but loses “Do not.” A keyword match or plausible sentence cannot substitute for checking meaning. When the answer controls a consequential action, require a reviewer to inspect the corresponding recording as well as the caption text.
Normalize Timestamps With an Explicit Contract
Convert declared units and timing semantics; never guess them from the magnitude of a number. A value of 12 could mean seconds, milliseconds or a frame count. A second field might be duration or end time. The collector’s documented format must establish the interpretation.
For a specific browser contract, MDN documents TextTrackCue.startTime as a number in seconds. That establishes the unit of that property, not the unit of every transcript API or subtitle export. Our adapter therefore requires an explicit unit on each authored input.
This self-contained Node.js script handles only two declared units, seconds and milliseconds, and exactly one of duration or end. Save it as transcript-fixture.mjs. It reads no files, calls no service and performs no video acquisition.
import assert from 'node:assert/strict';
function normalize(segment) {
if (!['seconds', 'milliseconds'].includes(segment.unit)) {
throw new Error('Unknown time unit');
}
const hasDuration = segment.duration !== undefined;
const hasEnd = segment.end !== undefined;
if (hasDuration === hasEnd) throw new Error('Declare duration OR end');
const values = [segment.start, hasDuration ? segment.duration : segment.end];
if (!values.every(value => typeof value === 'number' && Number.isFinite(value))) {
throw new Error('Timing values must be finite numbers');
}
const scale = segment.unit === 'milliseconds' ? 1000 : 1;
const start = segment.start / scale;
const end = hasDuration ? start + segment.duration / scale : segment.end / scale;
if (!Number.isFinite(end) || start < 0 || end <= start) {
throw new Error('Invalid interval');
}
if (typeof segment.text !== 'string' || !segment.text.trim()) {
throw new Error('Missing text');
}
return { id: segment.id, startSeconds: start, endSeconds: end,
text: segment.text.trim() };
}
const input = [
{ id: 'seg-01', unit: 'milliseconds', start: 12500, duration: 4000,
text: 'Do not enable exports before the workspace review.' },
{ id: 'seg-02', unit: 'seconds', start: 16.5, end: 22,
text: 'After approval, the administrator can enable exports.' },
{ id: 'seg-03', unit: 'seconds', start: 22, duration: 3,
text: 'Keep the review record with the workspace configuration.' }
];
const output = input.map(normalize);
assert.deepEqual(output.map(s => [s.startSeconds, s.endSeconds]),
[[12.5, 16.5], [16.5, 22], [22, 25]]);
assert.throws(() => normalize({ ...input[0], unit: 'unknown' }));
assert.throws(() => normalize({ ...input[1], duration: 1 }));
assert.throws(() => normalize({ ...input[2], duration: -1 }));
assert.throws(() => normalize({ ...input[0], start: '12500' }));
assert.throws(() => normalize({ ...input[2], start: Number.MAX_VALUE,
duration: Number.MAX_VALUE }));
console.log(JSON.stringify(output, null, 2));
The numeric-string rejection is deliberate: parsing belongs in a documented adapter, where an invalid value can be reported instead of silently coerced. The computed end must also remain finite: two finite inputs can overflow when added. This code is a normalization example, not a complete WebVTT parser. It does not handle clock strings, frame rates, cue settings, speaker markup or all subtitle formats.
The MDN example for the same cue interface creates a cue starting at 0.1 second and ending at 0.9 seconds. Its constructor example uses an end time; this does not mean a separate provider’s duration field should be read as an end time. The distinction is why the script refuses an input that declares both.
Do not clamp an invalid end time to make a row pass. Preserve the failure and inspect the contract. If the actual video duration is unknown, record that the upper-bound check was not performed. An interval can be internally valid while still pointing beyond the real recording.
Keep Context When Turning Segments Into Search Documents
Index passages that preserve the relationship between a statement and its condition. Caption boundaries often reflect display timing rather than complete ideas. A retrieval chunk may therefore need more than one segment.
For our fixture, a useful search document includes segments 01–02, the authored lesson title, track ID and combined range 12.5–22 seconds. It should also retain the original segment IDs so the interface can display exactly which words came from which interval. Do not replace the original timing with a guessed sentence-level alignment.
Normalize whitespace cautiously. Removing duplicate spaces may help search, but removing punctuation, speaker markers or short words can change meaning. Keep negations, quantities and qualifiers intact. If you translate or summarize, store that as a derived representation with a link to the original track, not as a verbatim transcript.
Repeated caption windows can produce repeated text. Deduplication should follow the actual format’s behavior and preserve timing evidence. Two speakers repeating the same sentence is different from an exporter duplicating the same cue. Do not delete every repeated string across a video merely because it appears redundant.
The website content for RAG workflow explains similar acceptance boundaries for web documents. The shared principle is preserving the evidence needed for an answer. The adapters differ: website section IDs are not video timestamps, and a caption track is not a complete record of visual information.
Validate Search Results and Source Navigation Separately
A correct answer, a relevant passage and a working source link are three separate checks. Test each one explicitly on the selected collection.
For a real video, navigate to the cited time and confirm the relevant passage can be found. A platform link may seek near a requested time rather than reproduce an exact frame. Keep the precise normalized range in your own record even if the player URL accepts only a whole-second offset. The synthetic fixture has no live video, so it intentionally provides no fabricated playback link.
Create an evaluation worksheet with the question, expected segment range, acceptable answer, prohibited inference, actual retrieval result and reviewer decision. Include a question whose answer is absent, a question spanning adjacent segments and one where an incorrect language track would mislead the reader.
| Check | Pass condition for the authored example | Failure example |
|---|---|---|
| Source identity | fixture-export-training, track fixture-en-v1 | Another lesson with similar wording |
| Evidence retrieval | Segments 01–02 available together | Segment 02 returned without the prerequisite |
| Meaning | Review precedes enabling exports | Negation omitted or order reversed |
| Citation | Track and 12.5–22 second range retained | Source title without a segment locator |
| Insufficient evidence | Missing prerequisite produces a review state | Model fills the gap with an invented rule |
This table states expected behavior; it does not report a measured system success rate. To publish one, run a declared evaluation set, retain actual outputs and include failed or unavailable cases in the appropriate denominator. A small handpicked demo should not become a universal accuracy claim.
Handle Missing Tracks and Conflicting Versions
Unavailable captions, a wrong-language track and failed collection must remain distinguishable. They require different follow-up and should not all produce an empty accepted transcript.
A missing required track may call for a permitted human transcription or an approved speech-to-text route. That is a separate workflow with its own cost, access conditions and quality checks. It is not a guaranteed free fallback, and text retrieval is not equivalent to transcribing audio.
If a video has both original and translated captions, preserve their relationship and allow the answer to identify which was used. Do not attribute translated wording as the speaker’s exact original phrasing. Where tracks disagree on an important number, name or negation, escalate rather than silently selecting the more convenient version.
Store the track version or a content fingerprint so later changes can be detected. If the source video is edited, previous timestamps may no longer locate the same material. Mark affected answers for revalidation instead of assuming a stable URL means stable content.
An internal library also needs removal and permission-change handling. If access is withdrawn, update retrieval eligibility and any derivative answer cache according to the owner’s policy. A user should not receive restricted content merely because it was indexed while they had different permissions.
Choose a Collection Tool by the Accepted Output
Evaluate the actual input, output and failure contract before comparing tool costs. Ask whether the permitted route can provide the needed track, timing semantics, language evidence, source identity and export format. Confirm these in current documentation and saved sample responses.
The YouTube Transcript Scraper guide provides an evaluation checklist for that candidate. For an Apify workflow, platform documentation is the reference for the configuration you actually use. Product names and a general platform integration do not establish access to every video or track.
Use the appropriate format documentation for the actual export, and inspect the source platform’s own caption documentation for acquisition permissions. The MDN cue-time reference above describes a browser representation; it does not grant download rights or an account entitlement. Acquisition permissions and current track availability need verification for your source and account.
The reviewed VidNotes course-transcript guide also identifies a practical boundary: visuals that an instructor does not narrate will not appear in a transcript. That vendor guide informed this article’s coverage; its accuracy, timing, pricing and learning-outcome claims are not evidence for this workflow. The two MDN links above refer to the same document, not independent sources.
Budget accepted transcripts and reviewed answers separately. If one transcript supports many questions, dividing all pipeline cost by a convenient answer count can hide review and maintenance work. Define the operating unit first, then use a ledger of collection, normalization, correction and storage costs. The AI agent cost worksheet helps make included components explicit without supplying a current quote for this task.
Deliver a Knowledge Base Someone Can Maintain
The handoff should explain how to reproduce an answer and how to stop using an unreliable source. Include the collection manifest, track schema, normalization rules, acceptance fixtures, evaluation questions, rejection states and update owner.
Keep the first release bounded enough for a reviewer to inspect its critical examples. Add sources when their access conditions and format are understood. More videos are not automatically more useful if users cannot distinguish a supported answer from a confident guess.
A maintenance review should look at failed retrieval questions, broken locators, unavailable tracks, permission changes and correction effort. Record what was observed instead of promising a fixed refresh cadence that nobody operates. Automated jobs need an actual configured owner and mechanism; a sentence in a guide does not create them.
If you need an ingestion implementation, request a web scraping automation scope with a permitted sample and required acceptance conditions. If your main problem is evaluating the resulting marketing knowledge workflow, the AI marketing audit offers a separate availability inquiry. Neither link implies this synthetic example has been tested on your collection.
Frequently Asked Questions
Can I build a knowledge base from plain transcript text?
Yes, but source navigation is limited if timing information is absent. Preserve the video and track identity, label the missing timing and avoid invented timestamps. Decide whether that limitation is acceptable for your task.
Are manually created captions always more accurate?
No accuracy assumption replaces checking the needed passage. Preserve the evidenced caption origin and inspect names, numbers and negations. If the origin is unknown, say so rather than guessing from writing quality.
Should I convert milliseconds to seconds automatically?
Only when the input contract establishes that the numbers are milliseconds. Divide by 1,000 in that adapter and keep the original values. Guessing the unit from a number’s size can create plausible but incorrect links.
What should I do with overlapping caption segments?
Inspect the format and playback context before treating overlap as an error. Simultaneous speakers or display behavior may be legitimate. Preserve the original intervals and flag cases that violate your documented adapter rules.
Can a transcript explain everything visible in a video?
No: captions may omit diagrams, gestures, screen values and other visual context. If the answer depends on those elements, inspect the recording through an authorized route or mark the evidence insufficient. Do not invent visual details from text.
Does the sample script download captions or prove a product works?
No: it normalizes three authored records and checks declared edge cases locally. Acquisition, track availability, playback navigation and retrieval accuracy need separate tests on a permitted real collection. No vendor performance result is claimed here.
Last verified: October 4, 2026. The MDN TextTrackCue startTime reference was read in full; its repeated links support different points from the same document, not independent studies. The reviewed VidNotes competitor guide informed topic coverage, but its accuracy, timing, pricing and learning-outcome claims were not adopted. The transcript and local normalization results are authored; current caption availability and product entitlements were not tested.
Web Scraping Automation
Apify scrapers, data extraction pipelines, and scheduled monitoring workflows.
View service detailsGoogle Preferred Sources
See more of my research in Google
Add Konabayev.com as a preferred source to find more fresh marketing and AI research in Google Search.
Have a relevant product or documented use case? View sponsorship formats and editorial requirements.


