📑 فهرست مطالب
🔍 تحلیل وضعیت فعلی
در نسخه فعلی XAudioFileContentExtractor، متد ConvertToWavAsync به صورت TODO رها شده و صرفاً استریم ورودی را بدون هیچ تبدیلی برمیگرداند:
// ❌ کد فعلی - ناقص
private async Task<Stream> ConvertToWavAsync(
Stream inputStream,
CancellationToken cancellationToken
)
{
// TODO: Complete this ...
return await Task.FromResult(inputStream);
}
⚠️ مشکل: مدل Whisper فقط فرمت WAV 16kHz Mono PCM 16-bit را قبول میکند. اگر فایل ورودی MP3، M4A، OGG یا هر فرمت دیگری باشد، Whisper خطا میدهد یا خروجی نادرست تولید میکند.
فرمتهای پشتیبانی شده و وضعیت فعلی:
| فرمت | MIME Type | وضعیت فعلی | نیاز به تبدیل |
|---|---|---|---|
| MP3 | audio/mpeg |
❌ بدون تبدیل | ✅ بله |
| WAV | audio/wav |
⚠️ ممکن است نیاز به resample داشته باشد | ⚠️ شاید |
| OGG | audio/ogg |
❌ بدون تبدیل | ✅ بله |
| M4A | audio/m4a |
❌ بدون تبدیل | ✅ بله |
| MP4 Audio | audio/mp4 |
❌ بدون تبدیل | ✅ بله |
| WebM Audio | audio/webm |
❌ بدون تبدیل | ✅ بله |
📦 گام ۱: نصب پکیجهای NuGet
برای تبدیل فرمتهای صوتی به WAV 16kHz، از کتابخانه NAudio استفاده میکنیم:
# در Package Manager Console: Install-Package NAudio # یا در .NET CLI: dotnet add package NAudio
💡 چرا NAudio؟
- پشتیبانی از MP3, WAV, AIFF و فرمتهای Windows Media
- قابلیت Resampling به هر نرخ نمونهبرداری
- تبدیل Stereo به Mono
- استفاده از Media Foundation ویندوز برای فرمتهای M4A, WMA, OGG
- بدون نیاز به نصب نرمافزار جانبی
🛠️ گام ۲: کد کامل XAudioFileContentExtractor
MODIFY
xAiApi/Providers/Extractors/XAudioFileContentExtractor.cs
using System;
using System.IO;
using System.Linq;
using NAudio.Wave;
using Whisper.net;
using System.Threading;
using xAiModels.Models;
using NAudio.MediaFoundation;
using System.Threading.Tasks;
using xAiApi.Interfaces.Extractors;
namespace xAiApi.Providers.Extractors
{
/// <summary>
/// Extracts text content from Audio files using Whisper ...
/// </summary>
public class XAudioFileContentExtractor : IXAudioFileContentExtractor
{
/// <summary>
/// Supported MIME Types ...
/// </summary>
private static readonly string[] SupportedMimeTypes =
[
"audio/mpeg",
"audio/mp3",
"audio/wav",
"audio/wave",
"audio/ogg",
"audio/m4a",
"audio/mp4",
"audio/webm"
];
private readonly string language;
private readonly string whisperModelPath;
public XAudioFileContentExtractor() : this(
language: "fa",
whisperModelPath: "Models/ggml-base.bin"
)
{ }
public XAudioFileContentExtractor(
string whisperModelPath = "Models/ggml-base.bin",
string language = "fa"
)
{
this.language = language;
this.whisperModelPath = whisperModelPath;
}
/// <summary>
/// Check if this extractor supports the specified MIME type ...
/// </summary>
public bool CanExtract(string mimeType)
{
return SupportedMimeTypes.Contains(
mimeType?.ToLowerInvariant() ?? string.Empty
);
}
/// <summary>
/// Extract text content from file stream ...
/// </summary>
public async Task<string> ExtractAsync(
Stream fileStream,
string mimeType,
CancellationToken cancellationToken = default
)
{
var result = await ExtractRichAsync(
fileStream,
"audio",
mimeType,
cancellationToken
);
return result.AudioTranscript;
}
/// <summary>
/// Extract content from stream as Rich Result ...
/// </summary>
public async Task<XFileExtractionResult> ExtractRichAsync(
Stream fileStream,
string fileName,
string mimeType,
CancellationToken cancellationToken = default
)
{
var result = new XFileExtractionResult
{
FileName = fileName,
MimeType = mimeType
};
try
{
// ۱. تبدیل فرمت صوتی به WAV 16kHz Mono
using var wavStream = await ConvertToWavAsync(
fileStream,
mimeType,
cancellationToken
);
// ۲. بررسی وجود مدل Whisper
if (!File.Exists(whisperModelPath))
{
result.ErrorMessage =
$"Whisper model not found at: {whisperModelPath}. " +
"Please download from https://huggingface.co/ggerganov/whisper.cpp/tree/main";
return result;
}
// ۳. انجام Speech-to-Text با Whisper
using var factory = WhisperFactory.FromPath(whisperModelPath);
using var processor = factory.CreateBuilder()
.WithLanguage(language)
.Build();
var segments = new System.Text.StringBuilder();
await foreach (var segment in processor.ProcessAsync(
wavStream,
cancellationToken))
{
segments.Append(segment.Text);
}
result.AudioTranscript = segments.ToString().Trim();
result.Text = result.AudioTranscript;
}
catch (OperationCanceledException)
{
throw;
}
catch (Exception ex)
{
result.ErrorMessage = $"Audio extraction failed: {ex.Message}";
}
return result;
}
/// <summary>
/// Convert any audio format to WAV 16kHz Mono 16-bit PCM
/// (required format for Whisper) ...
/// </summary>
private async Task<Stream> ConvertToWavAsync(
Stream inputStream,
string mimeType,
CancellationToken cancellationToken
)
{
return await Task.Run(() =>
{
// کپی به MemoryStream برای NAudio (نیاز به Seekable Stream)
var memoryStream = new MemoryStream();
inputStream.CopyTo(memoryStream);
memoryStream.Position = 0;
// فرمت هدف: 16kHz, 16-bit, Mono (الزامی برای Whisper)
var targetFormat = new WaveFormat(16000, 16, 1);
// خواندن فایل صوتی بر اساس فرمت
WaveStream reader = GetAudioReader(memoryStream, mimeType);
if (reader == null)
{
// Fallback: تلاش با MediaFoundationReader برای فرمتهای ناشناخته
try
{
memoryStream.Position = 0;
MediaFoundationApi.Startup();
reader = new MediaFoundationReader(memoryStream);
}
catch
{
memoryStream.Position = 0;
return memoryStream;
}
}
// بررسی آیا تبدیل لازم است یا خیر
var needsConversion =
reader.WaveFormat.SampleRate != 16000 ||
reader.WaveFormat.Channels != 1 ||
reader.WaveFormat.BitsPerSample != 16 ||
reader.WaveFormat.Encoding != WaveFormatEncoding.Pcm;
if (!needsConversion)
{
// فرمت صحیح است، نیازی به تبدیل نیست
reader.Dispose();
memoryStream.Position = 0;
return memoryStream;
}
// تبدیل فرمت با Resampler
var outputStream = new MemoryStream();
try
{
MediaFoundationApi.Startup();
using var resampler = new MediaFoundationResampler(
reader,
targetFormat
);
resampler.ResamplerQuality = 60;
WaveFileWriter.WriteWavFileToStream(outputStream, resampler);
outputStream.Position = 0;
reader.Dispose();
memoryStream.Dispose();
}
catch
{
// در صورت خطا در Resampler، استریم اصلی را برگردان
outputStream.Dispose();
reader.Dispose();
memoryStream.Position = 0;
return memoryStream;
}
return outputStream;
}, cancellationToken);
}
/// <summary>
/// Get appropriate WaveStream reader based on MIME type ...
/// </summary>
private WaveStream GetAudioReader(
MemoryStream stream,
string mimeType
)
{
try
{
var normalizedMime = mimeType?.ToLowerInvariant() ?? string.Empty;
switch (normalizedMime)
{
case "audio/mpeg":
case "audio/mp3":
return new Mp3FileReader(stream);
case "audio/wav":
case "audio/wave":
return new WaveFileReader(stream);
case "audio/ogg":
case "audio/m4a":
case "audio/mp4":
case "audio/webm":
// استفاده از MediaFoundation برای فرمتهای پیشرفته
MediaFoundationApi.Startup();
return new MediaFoundationReader(stream);
default:
return null;
}
}
catch
{
return null;
}
}
}
}
📥 گام ۳: آمادهسازی مدل Whisper
۳.۱. دانلود مدل:
مدلهای Whisper را از لینک زیر دانلود کنید:
https://huggingface.co/ggerganov/whisper.cpp/tree/main
۳.۲. مدلهای پیشنهادی:
| مدل | حجم | سرعت | دقت فارسی | کاربرد |
|---|---|---|---|---|
ggml-tiny.bin |
~75 MB | ⭐⭐⭐⭐⭐ | ⭐⭐ | تست و توسعه |
ggml-base.bin |
~142 MB | ⭐⭐⭐⭐ | ⭐⭐⭐ | استفاده عمومی (پیشفرض) |
ggml-small.bin |
~466 MB | ⭐⭐⭐ | ⭐⭐⭐⭐ | دقت بالاتر |
ggml-medium.bin |
~1.5 GB | ⭐⭐ | ⭐⭐⭐⭐⭐ | Production |
ggml-large-v3.bin |
~3 GB | ⭐ | ⭐⭐⭐⭐⭐ | حداکثر دقت |
۳.۳. ساختار پوشهها:
xAiApi/
└── Models/
├── ggml-base.bin ← پیشفرض
└── ggml-small.bin ← اختیاری (دقت بالاتر)
⚙️ گام ۴: پیکربندی appsettings.json
{
"AiApiConfiguration": {
"Models": [ ... ],
"Prompts": [ ... ],
"OCR": {
"DataPath": "tessdata",
"EngineMode": "LstmOnly",
"EnableOcrFallback": true,
"DefaultLanguage": "fas+eng"
},
"Audio": {
"WhisperModelPath": "Models/ggml-base.bin",
"Language": "fa"
}
}
}
💡 نکته: اگر از پیکربندی
Audio استفاده میکنید، باید Constructor کلاس XAudioFileContentExtractor را به صورت Factory در Startup.cs ثبت کنید تا مقادیر از IConfiguration خوانده شوند.
📋 خلاصه تغییرات
| فایل | تغییر | توضیح |
|---|---|---|
XAudioFileContentExtractor.cs |
MODIFY | پیادهسازی کامل ConvertToWavAsync با NAudio |
| NuGet Packages | INSTALL | نصب NAudio |
| Models/ | ADD | دانلود مدل ggml-base.bin از HuggingFace |
ویژگیهای کلیدی پیادهسازی:
🔄 تبدیل فرمت خودکار
- MP3, OGG, M4A, WebM → WAV 16kHz
- استفاده از MediaFoundationResampler
- تبدیل Stereo به Mono
⚡ بهینهسازی عملکرد
- بررسی نیاز به تبدیل قبل از پردازش
- رد کردن تبدیل اگر فرمت صحیح باشد
- اجرای async در Thread جداگانه
🛡️ مدیریت خطا
- بررسی وجود مدل Whisper
- Fallback در صورت خطای Resampler
- پشتیبانی از CancellationToken
🌐 پشتیبانی چند فرمتی
- Mp3FileReader برای MP3
- WaveFileReader برای WAV
- MediaFoundationReader برای M4A/OGG/WebM
✅ نتیجه نهایی:
سرویس XAudioFileContentExtractor اکنون به صورت کامل قادر است:
- هر فرمت صوتی رایج را به WAV 16kHz Mono تبدیل کند
- متن فارسی و انگلیسی را از فایل صوتی استخراج کند
- بدون نیاز به نرمافزار جانبی (ffmpeg و غیره) کار کند
- در محیط Production با اطمینان عمل کند