Files
xSaherelmWorkspace/Documents/Docs/Developer/15_AudioFileContentExtractor.html
T
2026-10-04 11:42:43 +03:30

701 lines
27 KiB
HTML
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
<!DOCTYPE html>
<html lang="fa" dir="rtl">
<head>
<meta charset="UTF-8">
<meta name="viewport" content="width=device-width, initial-scale=1.0">
<title>تکمیل سرویس XAudioFileContentExtractor - فن آوران ساحر علم</title>
<style>
:root {
--primary: #1e3a8a;
--secondary: #3b82f6;
--accent: #f59e0b;
--success: #10b981;
--danger: #ef4444;
--warning: #f97316;
--bg-light: #f8fafc;
--bg-code: #1e293b;
--text-dark: #0f172a;
--text-muted: #64748b;
--border: #e2e8f0;
}
* { box-sizing: border-box; margin: 0; padding: 0; }
body {
font-family: 'Tahoma', 'Segoe UI', sans-serif;
background: linear-gradient(135deg, #f8fafc 0%, #e0e7ff 100%);
color: var(--text-dark);
line-height: 1.8;
padding: 20px;
}
.container {
max-width: 1200px;
margin: 0 auto;
background: white;
border-radius: 16px;
box-shadow: 0 20px 60px rgba(0,0,0,0.1);
overflow: hidden;
}
.header {
background: linear-gradient(135deg, var(--primary) 0%, var(--secondary) 100%);
color: white;
padding: 40px;
text-align: center;
}
.header h1 { font-size: 2.1em; margin-bottom: 10px; }
.header .subtitle { font-size: 1.1em; opacity: 0.95; }
.meta-bar {
display: flex;
justify-content: space-between;
background: var(--bg-light);
padding: 15px 30px;
border-bottom: 2px solid var(--border);
flex-wrap: wrap;
gap: 15px;
}
.meta-item { display: flex; align-items: center; gap: 8px; font-size: 0.9em; color: var(--text-muted); }
.meta-item strong { color: var(--primary); }
.content { padding: 40px; }
.section {
margin-bottom: 35px;
padding: 25px;
background: var(--bg-light);
border-radius: 12px;
border-right: 5px solid var(--secondary);
}
.section h2 {
color: var(--primary);
font-size: 1.5em;
margin-bottom: 20px;
padding-bottom: 10px;
border-bottom: 2px solid var(--border);
}
.section h3 { color: var(--secondary); font-size: 1.2em; margin: 20px 0 12px; }
pre {
background: var(--bg-code);
color: #e2e8f0;
padding: 18px;
border-radius: 8px;
overflow-x: auto;
direction: ltr;
text-align: left;
font-family: 'Consolas', monospace;
font-size: 0.85em;
margin: 15px 0;
border-right: 4px solid var(--accent);
}
code {
background: #fef3c7;
color: #92400e;
padding: 2px 8px;
border-radius: 4px;
font-family: 'Consolas', monospace;
font-size: 0.9em;
direction: ltr;
display: inline-block;
}
table {
width: 100%;
border-collapse: collapse;
margin: 15px 0;
background: white;
border-radius: 8px;
overflow: hidden;
}
th { background: var(--primary); color: white; padding: 12px; text-align: right; }
td { padding: 12px; border-bottom: 1px solid var(--border); }
tr:hover { background: var(--bg-light); }
.alert { padding: 15px 20px; border-radius: 8px; margin: 15px 0; border-right: 4px solid; }
.alert-info { background: #dbeafe; border-color: var(--secondary); color: #1e40af; }
.alert-success { background: #d1fae5; border-color: var(--success); color: #065f46; }
.alert-warning { background: #fef3c7; border-color: var(--accent); color: #92400e; }
.alert-danger { background: #fee2e2; border-color: var(--danger); color: #991b1b; }
.footer { background: var(--primary); color: white; padding: 25px; text-align: center; }
.footer p { margin: 5px 0; }
.toc { background: white; padding: 20px; border-radius: 10px; margin-bottom: 25px; border: 2px solid var(--border); }
.toc h3 { color: var(--primary); margin-bottom: 15px; }
.toc ol { padding-right: 25px; }
.toc li { padding: 6px 0; }
.toc a { color: var(--secondary); text-decoration: none; }
.file-change { background: #f0f9ff; border-right: 4px solid var(--secondary); padding: 15px; margin: 10px 0; border-radius: 8px; }
.file-change .path { font-family: 'Consolas', monospace; color: var(--primary); font-weight: bold; direction: ltr; display: inline-block; }
.badge-modify { background: var(--warning); color: white; padding: 2px 8px; border-radius: 4px; font-size: 0.75em; margin-right: 8px; }
.badge-new { background: var(--success); color: white; padding: 2px 8px; border-radius: 4px; font-size: 0.75em; margin-right: 8px; }
.arch-grid { display: grid; grid-template-columns: repeat(auto-fit, minmax(280px, 1fr)); gap: 20px; margin: 20px 0; }
.arch-card { background: white; padding: 20px; border-radius: 10px; box-shadow: 0 4px 12px rgba(0,0,0,0.08); border-top: 4px solid var(--secondary); }
.arch-card h4 { color: var(--primary); margin-bottom: 12px; }
.arch-card ul { list-style: none; padding-right: 0; }
.arch-card li { padding: 6px 0; padding-right: 20px; position: relative; }
.arch-card li::before { content: '▸'; position: absolute; right: 0; color: var(--accent); font-weight: bold; }
</style>
</head>
<body>
<div class="container">
<div class="header">
<h1>🎙️ تکمیل سرویس XAudioFileContentExtractor</h1>
<div class="subtitle">پیاده‌سازی کامل تبدیل صوت به متن با Whisper و NAudio</div>
</div>
<div class="meta-bar">
<div class="meta-item">👨‍💻 <strong>توسعه‌دهنده:</strong> هادی خزاعی اصل</div>
<div class="meta-item">🏢 <strong>شرکت:</strong> فن آوران ساحر علم</div>
<div class="meta-item">📅 <strong>تاریخ:</strong> یکشنبه ۱۳ مهر ۱۴۰۵</div>
<div class="meta-item">📦 <strong>پروژه:</strong> xAiApi</div>
</div>
<div class="content">
<div class="toc">
<h3>📑 فهرست مطالب</h3>
<ol>
<li><a href="#analysis">تحلیل وضعیت فعلی</a></li>
<li><a href="#nuget">گام ۱: نصب پکیج‌های NuGet</a></li>
<li><a href="#code">گام ۲: کد کامل XAudioFileContentExtractor</a></li>
<li><a href="#whisper">گام ۳: آماده‌سازی مدل Whisper</a></li>
<li><a href="#config">گام ۴: پیکربندی appsettings.json</a></li>
<li><a href="#summary">خلاصه تغییرات</a></li>
</ol>
</div>
<!-- Section 1: Analysis -->
<div class="section" id="analysis">
<h2>🔍 تحلیل وضعیت فعلی</h2>
<p>در نسخه فعلی <code>XAudioFileContentExtractor</code>، متد <code>ConvertToWavAsync</code> به صورت <strong>TODO</strong> رها شده و صرفاً استریم ورودی را بدون هیچ تبدیلی برمی‌گرداند:</p>
<pre>// ❌ کد فعلی - ناقص
private async Task&lt;Stream&gt; ConvertToWavAsync(
Stream inputStream,
CancellationToken cancellationToken
)
{
// TODO: Complete this ...
return await Task.FromResult(inputStream);
}</pre>
<div class="alert alert-danger">
<strong>⚠️ مشکل:</strong> مدل Whisper فقط فرمت <strong>WAV 16kHz Mono PCM 16-bit</strong> را قبول می‌کند. اگر فایل ورودی MP3، M4A، OGG یا هر فرمت دیگری باشد، Whisper خطا می‌دهد یا خروجی نادرست تولید می‌کند.
</div>
<h3>فرمت‌های پشتیبانی شده و وضعیت فعلی:</h3>
<table>
<tr>
<th>فرمت</th>
<th>MIME Type</th>
<th>وضعیت فعلی</th>
<th>نیاز به تبدیل</th>
</tr>
<tr>
<td>MP3</td>
<td><code>audio/mpeg</code></td>
<td>❌ بدون تبدیل</td>
<td>✅ بله</td>
</tr>
<tr>
<td>WAV</td>
<td><code>audio/wav</code></td>
<td>⚠️ ممکن است نیاز به resample داشته باشد</td>
<td>⚠️ شاید</td>
</tr>
<tr>
<td>OGG</td>
<td><code>audio/ogg</code></td>
<td>❌ بدون تبدیل</td>
<td>✅ بله</td>
</tr>
<tr>
<td>M4A</td>
<td><code>audio/m4a</code></td>
<td>❌ بدون تبدیل</td>
<td>✅ بله</td>
</tr>
<tr>
<td>MP4 Audio</td>
<td><code>audio/mp4</code></td>
<td>❌ بدون تبدیل</td>
<td>✅ بله</td>
</tr>
<tr>
<td>WebM Audio</td>
<td><code>audio/webm</code></td>
<td>❌ بدون تبدیل</td>
<td>✅ بله</td>
</tr>
</table>
</div>
<!-- Section 2: NuGet -->
<div class="section" id="nuget">
<h2>📦 گام ۱: نصب پکیج‌های NuGet</h2>
<p>برای تبدیل فرمت‌های صوتی به WAV 16kHz، از کتابخانه <strong>NAudio</strong> استفاده می‌کنیم:</p>
<pre># در Package Manager Console:
Install-Package NAudio
# یا در .NET CLI:
dotnet add package NAudio</pre>
<div class="alert alert-info">
<strong>💡 چرا NAudio؟</strong>
<ul style="padding-right: 25px; margin-top: 10px;">
<li>پشتیبانی از MP3, WAV, AIFF و فرمت‌های Windows Media</li>
<li>قابلیت Resampling به هر نرخ نمونه‌برداری</li>
<li>تبدیل Stereo به Mono</li>
<li>استفاده از Media Foundation ویندوز برای فرمت‌های M4A, WMA, OGG</li>
<li>بدون نیاز به نصب نرم‌افزار جانبی</li>
</ul>
</div>
</div>
<!-- Section 3: Complete Code -->
<div class="section" id="code">
<h2>🛠️ گام ۲: کد کامل XAudioFileContentExtractor</h2>
<div class="file-change">
<span class="badge-modify">MODIFY</span>
<span class="path">xAiApi/Providers/Extractors/XAudioFileContentExtractor.cs</span>
</div>
<pre>using System;
using System.IO;
using System.Linq;
using NAudio.Wave;
using Whisper.net;
using System.Threading;
using xAiModels.Models;
using NAudio.MediaFoundation;
using System.Threading.Tasks;
using xAiApi.Interfaces.Extractors;
namespace xAiApi.Providers.Extractors
{
/// &lt;summary&gt;
/// Extracts text content from Audio files using Whisper ...
/// &lt;/summary&gt;
public class XAudioFileContentExtractor : IXAudioFileContentExtractor
{
/// &lt;summary&gt;
/// Supported MIME Types ...
/// &lt;/summary&gt;
private static readonly string[] SupportedMimeTypes =
[
"audio/mpeg",
"audio/mp3",
"audio/wav",
"audio/wave",
"audio/ogg",
"audio/m4a",
"audio/mp4",
"audio/webm"
];
private readonly string language;
private readonly string whisperModelPath;
public XAudioFileContentExtractor() : this(
language: "fa",
whisperModelPath: "Models/ggml-base.bin"
)
{ }
public XAudioFileContentExtractor(
string whisperModelPath = "Models/ggml-base.bin",
string language = "fa"
)
{
this.language = language;
this.whisperModelPath = whisperModelPath;
}
/// &lt;summary&gt;
/// Check if this extractor supports the specified MIME type ...
/// &lt;/summary&gt;
public bool CanExtract(string mimeType)
{
return SupportedMimeTypes.Contains(
mimeType?.ToLowerInvariant() ?? string.Empty
);
}
/// &lt;summary&gt;
/// Extract text content from file stream ...
/// &lt;/summary&gt;
public async Task&lt;string&gt; ExtractAsync(
Stream fileStream,
string mimeType,
CancellationToken cancellationToken = default
)
{
var result = await ExtractRichAsync(
fileStream,
"audio",
mimeType,
cancellationToken
);
return result.AudioTranscript;
}
/// &lt;summary&gt;
/// Extract content from stream as Rich Result ...
/// &lt;/summary&gt;
public async Task&lt;XFileExtractionResult&gt; ExtractRichAsync(
Stream fileStream,
string fileName,
string mimeType,
CancellationToken cancellationToken = default
)
{
var result = new XFileExtractionResult
{
FileName = fileName,
MimeType = mimeType
};
try
{
// ۱. تبدیل فرمت صوتی به WAV 16kHz Mono
using var wavStream = await ConvertToWavAsync(
fileStream,
mimeType,
cancellationToken
);
// ۲. بررسی وجود مدل Whisper
if (!File.Exists(whisperModelPath))
{
result.ErrorMessage =
$"Whisper model not found at: {whisperModelPath}. " +
"Please download from https://huggingface.co/ggerganov/whisper.cpp/tree/main";
return result;
}
// ۳. انجام Speech-to-Text با Whisper
using var factory = WhisperFactory.FromPath(whisperModelPath);
using var processor = factory.CreateBuilder()
.WithLanguage(language)
.Build();
var segments = new System.Text.StringBuilder();
await foreach (var segment in processor.ProcessAsync(
wavStream,
cancellationToken))
{
segments.Append(segment.Text);
}
result.AudioTranscript = segments.ToString().Trim();
result.Text = result.AudioTranscript;
}
catch (OperationCanceledException)
{
throw;
}
catch (Exception ex)
{
result.ErrorMessage = $"Audio extraction failed: {ex.Message}";
}
return result;
}
/// &lt;summary&gt;
/// Convert any audio format to WAV 16kHz Mono 16-bit PCM
/// (required format for Whisper) ...
/// &lt;/summary&gt;
private async Task&lt;Stream&gt; ConvertToWavAsync(
Stream inputStream,
string mimeType,
CancellationToken cancellationToken
)
{
return await Task.Run(() =&gt;
{
// کپی به MemoryStream برای NAudio (نیاز به Seekable Stream)
var memoryStream = new MemoryStream();
inputStream.CopyTo(memoryStream);
memoryStream.Position = 0;
// فرمت هدف: 16kHz, 16-bit, Mono (الزامی برای Whisper)
var targetFormat = new WaveFormat(16000, 16, 1);
// خواندن فایل صوتی بر اساس فرمت
WaveStream reader = GetAudioReader(memoryStream, mimeType);
if (reader == null)
{
// Fallback: تلاش با MediaFoundationReader برای فرمت‌های ناشناخته
try
{
memoryStream.Position = 0;
MediaFoundationApi.Startup();
reader = new MediaFoundationReader(memoryStream);
}
catch
{
memoryStream.Position = 0;
return memoryStream;
}
}
// بررسی آیا تبدیل لازم است یا خیر
var needsConversion =
reader.WaveFormat.SampleRate != 16000 ||
reader.WaveFormat.Channels != 1 ||
reader.WaveFormat.BitsPerSample != 16 ||
reader.WaveFormat.Encoding != WaveFormatEncoding.Pcm;
if (!needsConversion)
{
// فرمت صحیح است، نیازی به تبدیل نیست
reader.Dispose();
memoryStream.Position = 0;
return memoryStream;
}
// تبدیل فرمت با Resampler
var outputStream = new MemoryStream();
try
{
MediaFoundationApi.Startup();
using var resampler = new MediaFoundationResampler(
reader,
targetFormat
);
resampler.ResamplerQuality = 60;
WaveFileWriter.WriteWavFileToStream(outputStream, resampler);
outputStream.Position = 0;
reader.Dispose();
memoryStream.Dispose();
}
catch
{
// در صورت خطا در Resampler، استریم اصلی را برگردان
outputStream.Dispose();
reader.Dispose();
memoryStream.Position = 0;
return memoryStream;
}
return outputStream;
}, cancellationToken);
}
/// &lt;summary&gt;
/// Get appropriate WaveStream reader based on MIME type ...
/// &lt;/summary&gt;
private WaveStream GetAudioReader(
MemoryStream stream,
string mimeType
)
{
try
{
var normalizedMime = mimeType?.ToLowerInvariant() ?? string.Empty;
switch (normalizedMime)
{
case "audio/mpeg":
case "audio/mp3":
return new Mp3FileReader(stream);
case "audio/wav":
case "audio/wave":
return new WaveFileReader(stream);
case "audio/ogg":
case "audio/m4a":
case "audio/mp4":
case "audio/webm":
// استفاده از MediaFoundation برای فرمت‌های پیشرفته
MediaFoundationApi.Startup();
return new MediaFoundationReader(stream);
default:
return null;
}
}
catch
{
return null;
}
}
}
}</pre>
</div>
<!-- Section 4: Whisper Model -->
<div class="section" id="whisper">
<h2>📥 گام ۳: آماده‌سازی مدل Whisper</h2>
<h3>۳.۱. دانلود مدل:</h3>
<p>مدل‌های Whisper را از لینک زیر دانلود کنید:</p>
<p><a href="https://huggingface.co/ggerganov/whisper.cpp/tree/main" target="_blank">https://huggingface.co/ggerganov/whisper.cpp/tree/main</a></p>
<h3>۳.۲. مدل‌های پیشنهادی:</h3>
<table>
<tr>
<th>مدل</th>
<th>حجم</th>
<th>سرعت</th>
<th>دقت فارسی</th>
<th>کاربرد</th>
</tr>
<tr>
<td><code>ggml-tiny.bin</code></td>
<td>~75 MB</td>
<td>⭐⭐⭐⭐⭐</td>
<td>⭐⭐</td>
<td>تست و توسعه</td>
</tr>
<tr>
<td><code>ggml-base.bin</code></td>
<td>~142 MB</td>
<td>⭐⭐⭐⭐</td>
<td>⭐⭐⭐</td>
<td>استفاده عمومی (پیش‌فرض)</td>
</tr>
<tr>
<td><code>ggml-small.bin</code></td>
<td>~466 MB</td>
<td>⭐⭐⭐</td>
<td>⭐⭐⭐⭐</td>
<td>دقت بالاتر</td>
</tr>
<tr>
<td><code>ggml-medium.bin</code></td>
<td>~1.5 GB</td>
<td>⭐⭐</td>
<td>⭐⭐⭐⭐⭐</td>
<td>Production</td>
</tr>
<tr>
<td><code>ggml-large-v3.bin</code></td>
<td>~3 GB</td>
<td>⭐</td>
<td>⭐⭐⭐⭐⭐</td>
<td>حداکثر دقت</td>
</tr>
</table>
<h3>۳.۳. ساختار پوشه‌ها:</h3>
<pre>xAiApi/
└── Models/
├── ggml-base.bin ← پیش‌فرض
└── ggml-small.bin ← اختیاری (دقت بالاتر)</pre>
</div>
<!-- Section 5: Configuration -->
<div class="section" id="config">
<h2>⚙️ گام ۴: پیکربندی appsettings.json</h2>
<pre>{
"AiApiConfiguration": {
"Models": [ ... ],
"Prompts": [ ... ],
"OCR": {
"DataPath": "tessdata",
"EngineMode": "LstmOnly",
"EnableOcrFallback": true,
"DefaultLanguage": "fas+eng"
},
"Audio": {
"WhisperModelPath": "Models/ggml-base.bin",
"Language": "fa"
}
}
}</pre>
<div class="alert alert-info">
<strong>💡 نکته:</strong> اگر از پیکربندی <code>Audio</code> استفاده می‌کنید، باید Constructor کلاس <code>XAudioFileContentExtractor</code> را به صورت Factory در <code>Startup.cs</code> ثبت کنید تا مقادیر از <code>IConfiguration</code> خوانده شوند.
</div>
</div>
<!-- Section 6: Summary -->
<div class="section" id="summary">
<h2>📋 خلاصه تغییرات</h2>
<table>
<tr>
<th>فایل</th>
<th>تغییر</th>
<th>توضیح</th>
</tr>
<tr>
<td><code>XAudioFileContentExtractor.cs</code></td>
<td><span class="badge-modify">MODIFY</span></td>
<td>پیاده‌سازی کامل <code>ConvertToWavAsync</code> با NAudio</td>
</tr>
<tr>
<td>NuGet Packages</td>
<td><span class="badge-new">INSTALL</span></td>
<td>نصب <code>NAudio</code></td>
</tr>
<tr>
<td>Models/</td>
<td><span class="badge-new">ADD</span></td>
<td>دانلود مدل <code>ggml-base.bin</code> از HuggingFace</td>
</tr>
</table>
<h3>ویژگی‌های کلیدی پیاده‌سازی:</h3>
<div class="arch-grid">
<div class="arch-card">
<h4>🔄 تبدیل فرمت خودکار</h4>
<ul>
<li>MP3, OGG, M4A, WebM → WAV 16kHz</li>
<li>استفاده از MediaFoundationResampler</li>
<li>تبدیل Stereo به Mono</li>
</ul>
</div>
<div class="arch-card">
<h4>⚡ بهینه‌سازی عملکرد</h4>
<ul>
<li>بررسی نیاز به تبدیل قبل از پردازش</li>
<li>رد کردن تبدیل اگر فرمت صحیح باشد</li>
<li>اجرای async در Thread جداگانه</li>
</ul>
</div>
<div class="arch-card">
<h4>🛡️ مدیریت خطا</h4>
<ul>
<li>بررسی وجود مدل Whisper</li>
<li>Fallback در صورت خطای Resampler</li>
<li>پشتیبانی از CancellationToken</li>
</ul>
</div>
<div class="arch-card">
<h4>🌐 پشتیبانی چند فرمتی</h4>
<ul>
<li>Mp3FileReader برای MP3</li>
<li>WaveFileReader برای WAV</li>
<li>MediaFoundationReader برای M4A/OGG/WebM</li>
</ul>
</div>
</div>
<div class="alert alert-success">
<strong>✅ نتیجه نهایی:</strong>
<p>سرویس <code>XAudioFileContentExtractor</code> اکنون به صورت کامل قادر است:</p>
<ul style="padding-right: 25px; margin-top: 10px;">
<li>هر فرمت صوتی رایج را به WAV 16kHz Mono تبدیل کند</li>
<li>متن فارسی و انگلیسی را از فایل صوتی استخراج کند</li>
<li>بدون نیاز به نرم‌افزار جانبی (ffmpeg و غیره) کار کند</li>
<li>در محیط Production با اطمینان عمل کند</li>
</ul>
</div>
</div>
</div>
<div class="footer">
<p><strong>👨‍💻 توسعه‌دهنده:</strong> هادی خزاعی اصل</p>
<p><strong>🏢 شرکت:</strong> فن آوران ساحر علم</p>
<p><strong>📅 تاریخ:</strong> یکشنبه ۱۳ مهر ۱۴۰۵</p>
<p style="margin-top: 15px; opacity: 0.8; font-size: 0.9em;">
🎙️ تکمیل سرویس XAudioFileContentExtractor - تمامی حقوق محفوظ است
</p>
</div>
</div>
</body>
</html>