This commit is contained in:
2026-10-04 11:42:43 +03:30
parent c81bc0e923
commit d7054eeb73
5 changed files with 8705 additions and 1 deletions
@@ -0,0 +1,701 @@
<!DOCTYPE html>
<html lang="fa" dir="rtl">
<head>
<meta charset="UTF-8">
<meta name="viewport" content="width=device-width, initial-scale=1.0">
<title>تکمیل سرویس XAudioFileContentExtractor - فن آوران ساحر علم</title>
<style>
:root {
--primary: #1e3a8a;
--secondary: #3b82f6;
--accent: #f59e0b;
--success: #10b981;
--danger: #ef4444;
--warning: #f97316;
--bg-light: #f8fafc;
--bg-code: #1e293b;
--text-dark: #0f172a;
--text-muted: #64748b;
--border: #e2e8f0;
}
* { box-sizing: border-box; margin: 0; padding: 0; }
body {
font-family: 'Tahoma', 'Segoe UI', sans-serif;
background: linear-gradient(135deg, #f8fafc 0%, #e0e7ff 100%);
color: var(--text-dark);
line-height: 1.8;
padding: 20px;
}
.container {
max-width: 1200px;
margin: 0 auto;
background: white;
border-radius: 16px;
box-shadow: 0 20px 60px rgba(0,0,0,0.1);
overflow: hidden;
}
.header {
background: linear-gradient(135deg, var(--primary) 0%, var(--secondary) 100%);
color: white;
padding: 40px;
text-align: center;
}
.header h1 { font-size: 2.1em; margin-bottom: 10px; }
.header .subtitle { font-size: 1.1em; opacity: 0.95; }
.meta-bar {
display: flex;
justify-content: space-between;
background: var(--bg-light);
padding: 15px 30px;
border-bottom: 2px solid var(--border);
flex-wrap: wrap;
gap: 15px;
}
.meta-item { display: flex; align-items: center; gap: 8px; font-size: 0.9em; color: var(--text-muted); }
.meta-item strong { color: var(--primary); }
.content { padding: 40px; }
.section {
margin-bottom: 35px;
padding: 25px;
background: var(--bg-light);
border-radius: 12px;
border-right: 5px solid var(--secondary);
}
.section h2 {
color: var(--primary);
font-size: 1.5em;
margin-bottom: 20px;
padding-bottom: 10px;
border-bottom: 2px solid var(--border);
}
.section h3 { color: var(--secondary); font-size: 1.2em; margin: 20px 0 12px; }
pre {
background: var(--bg-code);
color: #e2e8f0;
padding: 18px;
border-radius: 8px;
overflow-x: auto;
direction: ltr;
text-align: left;
font-family: 'Consolas', monospace;
font-size: 0.85em;
margin: 15px 0;
border-right: 4px solid var(--accent);
}
code {
background: #fef3c7;
color: #92400e;
padding: 2px 8px;
border-radius: 4px;
font-family: 'Consolas', monospace;
font-size: 0.9em;
direction: ltr;
display: inline-block;
}
table {
width: 100%;
border-collapse: collapse;
margin: 15px 0;
background: white;
border-radius: 8px;
overflow: hidden;
}
th { background: var(--primary); color: white; padding: 12px; text-align: right; }
td { padding: 12px; border-bottom: 1px solid var(--border); }
tr:hover { background: var(--bg-light); }
.alert { padding: 15px 20px; border-radius: 8px; margin: 15px 0; border-right: 4px solid; }
.alert-info { background: #dbeafe; border-color: var(--secondary); color: #1e40af; }
.alert-success { background: #d1fae5; border-color: var(--success); color: #065f46; }
.alert-warning { background: #fef3c7; border-color: var(--accent); color: #92400e; }
.alert-danger { background: #fee2e2; border-color: var(--danger); color: #991b1b; }
.footer { background: var(--primary); color: white; padding: 25px; text-align: center; }
.footer p { margin: 5px 0; }
.toc { background: white; padding: 20px; border-radius: 10px; margin-bottom: 25px; border: 2px solid var(--border); }
.toc h3 { color: var(--primary); margin-bottom: 15px; }
.toc ol { padding-right: 25px; }
.toc li { padding: 6px 0; }
.toc a { color: var(--secondary); text-decoration: none; }
.file-change { background: #f0f9ff; border-right: 4px solid var(--secondary); padding: 15px; margin: 10px 0; border-radius: 8px; }
.file-change .path { font-family: 'Consolas', monospace; color: var(--primary); font-weight: bold; direction: ltr; display: inline-block; }
.badge-modify { background: var(--warning); color: white; padding: 2px 8px; border-radius: 4px; font-size: 0.75em; margin-right: 8px; }
.badge-new { background: var(--success); color: white; padding: 2px 8px; border-radius: 4px; font-size: 0.75em; margin-right: 8px; }
.arch-grid { display: grid; grid-template-columns: repeat(auto-fit, minmax(280px, 1fr)); gap: 20px; margin: 20px 0; }
.arch-card { background: white; padding: 20px; border-radius: 10px; box-shadow: 0 4px 12px rgba(0,0,0,0.08); border-top: 4px solid var(--secondary); }
.arch-card h4 { color: var(--primary); margin-bottom: 12px; }
.arch-card ul { list-style: none; padding-right: 0; }
.arch-card li { padding: 6px 0; padding-right: 20px; position: relative; }
.arch-card li::before { content: '▸'; position: absolute; right: 0; color: var(--accent); font-weight: bold; }
</style>
</head>
<body>
<div class="container">
<div class="header">
<h1>🎙️ تکمیل سرویس XAudioFileContentExtractor</h1>
<div class="subtitle">پیاده‌سازی کامل تبدیل صوت به متن با Whisper و NAudio</div>
</div>
<div class="meta-bar">
<div class="meta-item">👨‍💻 <strong>توسعه‌دهنده:</strong> هادی خزاعی اصل</div>
<div class="meta-item">🏢 <strong>شرکت:</strong> فن آوران ساحر علم</div>
<div class="meta-item">📅 <strong>تاریخ:</strong> یکشنبه ۱۳ مهر ۱۴۰۵</div>
<div class="meta-item">📦 <strong>پروژه:</strong> xAiApi</div>
</div>
<div class="content">
<div class="toc">
<h3>📑 فهرست مطالب</h3>
<ol>
<li><a href="#analysis">تحلیل وضعیت فعلی</a></li>
<li><a href="#nuget">گام ۱: نصب پکیج‌های NuGet</a></li>
<li><a href="#code">گام ۲: کد کامل XAudioFileContentExtractor</a></li>
<li><a href="#whisper">گام ۳: آماده‌سازی مدل Whisper</a></li>
<li><a href="#config">گام ۴: پیکربندی appsettings.json</a></li>
<li><a href="#summary">خلاصه تغییرات</a></li>
</ol>
</div>
<!-- Section 1: Analysis -->
<div class="section" id="analysis">
<h2>🔍 تحلیل وضعیت فعلی</h2>
<p>در نسخه فعلی <code>XAudioFileContentExtractor</code>، متد <code>ConvertToWavAsync</code> به صورت <strong>TODO</strong> رها شده و صرفاً استریم ورودی را بدون هیچ تبدیلی برمی‌گرداند:</p>
<pre>// ❌ کد فعلی - ناقص
private async Task&lt;Stream&gt; ConvertToWavAsync(
Stream inputStream,
CancellationToken cancellationToken
)
{
// TODO: Complete this ...
return await Task.FromResult(inputStream);
}</pre>
<div class="alert alert-danger">
<strong>⚠️ مشکل:</strong> مدل Whisper فقط فرمت <strong>WAV 16kHz Mono PCM 16-bit</strong> را قبول می‌کند. اگر فایل ورودی MP3، M4A، OGG یا هر فرمت دیگری باشد، Whisper خطا می‌دهد یا خروجی نادرست تولید می‌کند.
</div>
<h3>فرمت‌های پشتیبانی شده و وضعیت فعلی:</h3>
<table>
<tr>
<th>فرمت</th>
<th>MIME Type</th>
<th>وضعیت فعلی</th>
<th>نیاز به تبدیل</th>
</tr>
<tr>
<td>MP3</td>
<td><code>audio/mpeg</code></td>
<td>❌ بدون تبدیل</td>
<td>✅ بله</td>
</tr>
<tr>
<td>WAV</td>
<td><code>audio/wav</code></td>
<td>⚠️ ممکن است نیاز به resample داشته باشد</td>
<td>⚠️ شاید</td>
</tr>
<tr>
<td>OGG</td>
<td><code>audio/ogg</code></td>
<td>❌ بدون تبدیل</td>
<td>✅ بله</td>
</tr>
<tr>
<td>M4A</td>
<td><code>audio/m4a</code></td>
<td>❌ بدون تبدیل</td>
<td>✅ بله</td>
</tr>
<tr>
<td>MP4 Audio</td>
<td><code>audio/mp4</code></td>
<td>❌ بدون تبدیل</td>
<td>✅ بله</td>
</tr>
<tr>
<td>WebM Audio</td>
<td><code>audio/webm</code></td>
<td>❌ بدون تبدیل</td>
<td>✅ بله</td>
</tr>
</table>
</div>
<!-- Section 2: NuGet -->
<div class="section" id="nuget">
<h2>📦 گام ۱: نصب پکیج‌های NuGet</h2>
<p>برای تبدیل فرمت‌های صوتی به WAV 16kHz، از کتابخانه <strong>NAudio</strong> استفاده می‌کنیم:</p>
<pre># در Package Manager Console:
Install-Package NAudio
# یا در .NET CLI:
dotnet add package NAudio</pre>
<div class="alert alert-info">
<strong>💡 چرا NAudio؟</strong>
<ul style="padding-right: 25px; margin-top: 10px;">
<li>پشتیبانی از MP3, WAV, AIFF و فرمت‌های Windows Media</li>
<li>قابلیت Resampling به هر نرخ نمونه‌برداری</li>
<li>تبدیل Stereo به Mono</li>
<li>استفاده از Media Foundation ویندوز برای فرمت‌های M4A, WMA, OGG</li>
<li>بدون نیاز به نصب نرم‌افزار جانبی</li>
</ul>
</div>
</div>
<!-- Section 3: Complete Code -->
<div class="section" id="code">
<h2>🛠️ گام ۲: کد کامل XAudioFileContentExtractor</h2>
<div class="file-change">
<span class="badge-modify">MODIFY</span>
<span class="path">xAiApi/Providers/Extractors/XAudioFileContentExtractor.cs</span>
</div>
<pre>using System;
using System.IO;
using System.Linq;
using NAudio.Wave;
using Whisper.net;
using System.Threading;
using xAiModels.Models;
using NAudio.MediaFoundation;
using System.Threading.Tasks;
using xAiApi.Interfaces.Extractors;
namespace xAiApi.Providers.Extractors
{
/// &lt;summary&gt;
/// Extracts text content from Audio files using Whisper ...
/// &lt;/summary&gt;
public class XAudioFileContentExtractor : IXAudioFileContentExtractor
{
/// &lt;summary&gt;
/// Supported MIME Types ...
/// &lt;/summary&gt;
private static readonly string[] SupportedMimeTypes =
[
"audio/mpeg",
"audio/mp3",
"audio/wav",
"audio/wave",
"audio/ogg",
"audio/m4a",
"audio/mp4",
"audio/webm"
];
private readonly string language;
private readonly string whisperModelPath;
public XAudioFileContentExtractor() : this(
language: "fa",
whisperModelPath: "Models/ggml-base.bin"
)
{ }
public XAudioFileContentExtractor(
string whisperModelPath = "Models/ggml-base.bin",
string language = "fa"
)
{
this.language = language;
this.whisperModelPath = whisperModelPath;
}
/// &lt;summary&gt;
/// Check if this extractor supports the specified MIME type ...
/// &lt;/summary&gt;
public bool CanExtract(string mimeType)
{
return SupportedMimeTypes.Contains(
mimeType?.ToLowerInvariant() ?? string.Empty
);
}
/// &lt;summary&gt;
/// Extract text content from file stream ...
/// &lt;/summary&gt;
public async Task&lt;string&gt; ExtractAsync(
Stream fileStream,
string mimeType,
CancellationToken cancellationToken = default
)
{
var result = await ExtractRichAsync(
fileStream,
"audio",
mimeType,
cancellationToken
);
return result.AudioTranscript;
}
/// &lt;summary&gt;
/// Extract content from stream as Rich Result ...
/// &lt;/summary&gt;
public async Task&lt;XFileExtractionResult&gt; ExtractRichAsync(
Stream fileStream,
string fileName,
string mimeType,
CancellationToken cancellationToken = default
)
{
var result = new XFileExtractionResult
{
FileName = fileName,
MimeType = mimeType
};
try
{
// ۱. تبدیل فرمت صوتی به WAV 16kHz Mono
using var wavStream = await ConvertToWavAsync(
fileStream,
mimeType,
cancellationToken
);
// ۲. بررسی وجود مدل Whisper
if (!File.Exists(whisperModelPath))
{
result.ErrorMessage =
$"Whisper model not found at: {whisperModelPath}. " +
"Please download from https://huggingface.co/ggerganov/whisper.cpp/tree/main";
return result;
}
// ۳. انجام Speech-to-Text با Whisper
using var factory = WhisperFactory.FromPath(whisperModelPath);
using var processor = factory.CreateBuilder()
.WithLanguage(language)
.Build();
var segments = new System.Text.StringBuilder();
await foreach (var segment in processor.ProcessAsync(
wavStream,
cancellationToken))
{
segments.Append(segment.Text);
}
result.AudioTranscript = segments.ToString().Trim();
result.Text = result.AudioTranscript;
}
catch (OperationCanceledException)
{
throw;
}
catch (Exception ex)
{
result.ErrorMessage = $"Audio extraction failed: {ex.Message}";
}
return result;
}
/// &lt;summary&gt;
/// Convert any audio format to WAV 16kHz Mono 16-bit PCM
/// (required format for Whisper) ...
/// &lt;/summary&gt;
private async Task&lt;Stream&gt; ConvertToWavAsync(
Stream inputStream,
string mimeType,
CancellationToken cancellationToken
)
{
return await Task.Run(() =&gt;
{
// کپی به MemoryStream برای NAudio (نیاز به Seekable Stream)
var memoryStream = new MemoryStream();
inputStream.CopyTo(memoryStream);
memoryStream.Position = 0;
// فرمت هدف: 16kHz, 16-bit, Mono (الزامی برای Whisper)
var targetFormat = new WaveFormat(16000, 16, 1);
// خواندن فایل صوتی بر اساس فرمت
WaveStream reader = GetAudioReader(memoryStream, mimeType);
if (reader == null)
{
// Fallback: تلاش با MediaFoundationReader برای فرمت‌های ناشناخته
try
{
memoryStream.Position = 0;
MediaFoundationApi.Startup();
reader = new MediaFoundationReader(memoryStream);
}
catch
{
memoryStream.Position = 0;
return memoryStream;
}
}
// بررسی آیا تبدیل لازم است یا خیر
var needsConversion =
reader.WaveFormat.SampleRate != 16000 ||
reader.WaveFormat.Channels != 1 ||
reader.WaveFormat.BitsPerSample != 16 ||
reader.WaveFormat.Encoding != WaveFormatEncoding.Pcm;
if (!needsConversion)
{
// فرمت صحیح است، نیازی به تبدیل نیست
reader.Dispose();
memoryStream.Position = 0;
return memoryStream;
}
// تبدیل فرمت با Resampler
var outputStream = new MemoryStream();
try
{
MediaFoundationApi.Startup();
using var resampler = new MediaFoundationResampler(
reader,
targetFormat
);
resampler.ResamplerQuality = 60;
WaveFileWriter.WriteWavFileToStream(outputStream, resampler);
outputStream.Position = 0;
reader.Dispose();
memoryStream.Dispose();
}
catch
{
// در صورت خطا در Resampler، استریم اصلی را برگردان
outputStream.Dispose();
reader.Dispose();
memoryStream.Position = 0;
return memoryStream;
}
return outputStream;
}, cancellationToken);
}
/// &lt;summary&gt;
/// Get appropriate WaveStream reader based on MIME type ...
/// &lt;/summary&gt;
private WaveStream GetAudioReader(
MemoryStream stream,
string mimeType
)
{
try
{
var normalizedMime = mimeType?.ToLowerInvariant() ?? string.Empty;
switch (normalizedMime)
{
case "audio/mpeg":
case "audio/mp3":
return new Mp3FileReader(stream);
case "audio/wav":
case "audio/wave":
return new WaveFileReader(stream);
case "audio/ogg":
case "audio/m4a":
case "audio/mp4":
case "audio/webm":
// استفاده از MediaFoundation برای فرمت‌های پیشرفته
MediaFoundationApi.Startup();
return new MediaFoundationReader(stream);
default:
return null;
}
}
catch
{
return null;
}
}
}
}</pre>
</div>
<!-- Section 4: Whisper Model -->
<div class="section" id="whisper">
<h2>📥 گام ۳: آماده‌سازی مدل Whisper</h2>
<h3>۳.۱. دانلود مدل:</h3>
<p>مدل‌های Whisper را از لینک زیر دانلود کنید:</p>
<p><a href="https://huggingface.co/ggerganov/whisper.cpp/tree/main" target="_blank">https://huggingface.co/ggerganov/whisper.cpp/tree/main</a></p>
<h3>۳.۲. مدل‌های پیشنهادی:</h3>
<table>
<tr>
<th>مدل</th>
<th>حجم</th>
<th>سرعت</th>
<th>دقت فارسی</th>
<th>کاربرد</th>
</tr>
<tr>
<td><code>ggml-tiny.bin</code></td>
<td>~75 MB</td>
<td>⭐⭐⭐⭐⭐</td>
<td>⭐⭐</td>
<td>تست و توسعه</td>
</tr>
<tr>
<td><code>ggml-base.bin</code></td>
<td>~142 MB</td>
<td>⭐⭐⭐⭐</td>
<td>⭐⭐⭐</td>
<td>استفاده عمومی (پیش‌فرض)</td>
</tr>
<tr>
<td><code>ggml-small.bin</code></td>
<td>~466 MB</td>
<td>⭐⭐⭐</td>
<td>⭐⭐⭐⭐</td>
<td>دقت بالاتر</td>
</tr>
<tr>
<td><code>ggml-medium.bin</code></td>
<td>~1.5 GB</td>
<td>⭐⭐</td>
<td>⭐⭐⭐⭐⭐</td>
<td>Production</td>
</tr>
<tr>
<td><code>ggml-large-v3.bin</code></td>
<td>~3 GB</td>
<td>⭐</td>
<td>⭐⭐⭐⭐⭐</td>
<td>حداکثر دقت</td>
</tr>
</table>
<h3>۳.۳. ساختار پوشه‌ها:</h3>
<pre>xAiApi/
└── Models/
├── ggml-base.bin ← پیش‌فرض
└── ggml-small.bin ← اختیاری (دقت بالاتر)</pre>
</div>
<!-- Section 5: Configuration -->
<div class="section" id="config">
<h2>⚙️ گام ۴: پیکربندی appsettings.json</h2>
<pre>{
"AiApiConfiguration": {
"Models": [ ... ],
"Prompts": [ ... ],
"OCR": {
"DataPath": "tessdata",
"EngineMode": "LstmOnly",
"EnableOcrFallback": true,
"DefaultLanguage": "fas+eng"
},
"Audio": {
"WhisperModelPath": "Models/ggml-base.bin",
"Language": "fa"
}
}
}</pre>
<div class="alert alert-info">
<strong>💡 نکته:</strong> اگر از پیکربندی <code>Audio</code> استفاده می‌کنید، باید Constructor کلاس <code>XAudioFileContentExtractor</code> را به صورت Factory در <code>Startup.cs</code> ثبت کنید تا مقادیر از <code>IConfiguration</code> خوانده شوند.
</div>
</div>
<!-- Section 6: Summary -->
<div class="section" id="summary">
<h2>📋 خلاصه تغییرات</h2>
<table>
<tr>
<th>فایل</th>
<th>تغییر</th>
<th>توضیح</th>
</tr>
<tr>
<td><code>XAudioFileContentExtractor.cs</code></td>
<td><span class="badge-modify">MODIFY</span></td>
<td>پیاده‌سازی کامل <code>ConvertToWavAsync</code> با NAudio</td>
</tr>
<tr>
<td>NuGet Packages</td>
<td><span class="badge-new">INSTALL</span></td>
<td>نصب <code>NAudio</code></td>
</tr>
<tr>
<td>Models/</td>
<td><span class="badge-new">ADD</span></td>
<td>دانلود مدل <code>ggml-base.bin</code> از HuggingFace</td>
</tr>
</table>
<h3>ویژگی‌های کلیدی پیاده‌سازی:</h3>
<div class="arch-grid">
<div class="arch-card">
<h4>🔄 تبدیل فرمت خودکار</h4>
<ul>
<li>MP3, OGG, M4A, WebM → WAV 16kHz</li>
<li>استفاده از MediaFoundationResampler</li>
<li>تبدیل Stereo به Mono</li>
</ul>
</div>
<div class="arch-card">
<h4>⚡ بهینه‌سازی عملکرد</h4>
<ul>
<li>بررسی نیاز به تبدیل قبل از پردازش</li>
<li>رد کردن تبدیل اگر فرمت صحیح باشد</li>
<li>اجرای async در Thread جداگانه</li>
</ul>
</div>
<div class="arch-card">
<h4>🛡️ مدیریت خطا</h4>
<ul>
<li>بررسی وجود مدل Whisper</li>
<li>Fallback در صورت خطای Resampler</li>
<li>پشتیبانی از CancellationToken</li>
</ul>
</div>
<div class="arch-card">
<h4>🌐 پشتیبانی چند فرمتی</h4>
<ul>
<li>Mp3FileReader برای MP3</li>
<li>WaveFileReader برای WAV</li>
<li>MediaFoundationReader برای M4A/OGG/WebM</li>
</ul>
</div>
</div>
<div class="alert alert-success">
<strong>✅ نتیجه نهایی:</strong>
<p>سرویس <code>XAudioFileContentExtractor</code> اکنون به صورت کامل قادر است:</p>
<ul style="padding-right: 25px; margin-top: 10px;">
<li>هر فرمت صوتی رایج را به WAV 16kHz Mono تبدیل کند</li>
<li>متن فارسی و انگلیسی را از فایل صوتی استخراج کند</li>
<li>بدون نیاز به نرم‌افزار جانبی (ffmpeg و غیره) کار کند</li>
<li>در محیط Production با اطمینان عمل کند</li>
</ul>
</div>
</div>
</div>
<div class="footer">
<p><strong>👨‍💻 توسعه‌دهنده:</strong> هادی خزاعی اصل</p>
<p><strong>🏢 شرکت:</strong> فن آوران ساحر علم</p>
<p><strong>📅 تاریخ:</strong> یکشنبه ۱۳ مهر ۱۴۰۵</p>
<p style="margin-top: 15px; opacity: 0.8; font-size: 0.9em;">
🎙️ تکمیل سرویس XAudioFileContentExtractor - تمامی حقوق محفوظ است
</p>
</div>
</div>
</body>
</html>