feat: centralize common English words for language detection
Move COMMON_ENGLISH_WORDS set from youtube_transcript_downloader.py to language_codes.py as a shared constant. This improves code organization and reusability, allowing other modules to reference the same word list for English language detection.
This commit is contained in:
@@ -19,12 +19,14 @@ def extract_video_id(url):
|
||||
patterns = [
|
||||
r'(?:v=|\/)([0-9A-Za-z_-]{11}).*',
|
||||
r'youtu\.be\/([0-9A-Za-z_-]{11})',
|
||||
r'([0-9A-Za-z_-]{11})'
|
||||
]
|
||||
|
||||
for pattern in patterns:
|
||||
match = re.search(pattern, url)
|
||||
if match:
|
||||
return match.group(1)
|
||||
|
||||
return None
|
||||
|
||||
|
||||
|
||||
Reference in New Issue
Block a user