638 lines
30 KiB
HTML
638 lines
30 KiB
HTML
<!DOCTYPE html>
|
||
<html lang="fa" dir="rtl">
|
||
<head>
|
||
<meta charset="UTF-8">
|
||
<meta name="viewport" content="width=device-width, initial-scale=1.0">
|
||
<title>استخراج محتوای جهانی با Qwen3.5-9B Vision - فن آوران ساحر علم</title>
|
||
<style>
|
||
:root {
|
||
--primary: #1e3a8a;
|
||
--secondary: #3b82f6;
|
||
--accent: #f59e0b;
|
||
--success: #10b981;
|
||
--danger: #ef4444;
|
||
--warning: #f97316;
|
||
--bg-light: #f8fafc;
|
||
--bg-code: #1e293b;
|
||
--text-dark: #0f172a;
|
||
--text-muted: #64748b;
|
||
--border: #e2e8f0;
|
||
}
|
||
* { box-sizing: border-box; margin: 0; padding: 0; }
|
||
body {
|
||
font-family: 'Tahoma', 'Segoe UI', sans-serif;
|
||
background: linear-gradient(135deg, #f8fafc 0%, #e0e7ff 100%);
|
||
color: var(--text-dark);
|
||
line-height: 1.8;
|
||
padding: 20px;
|
||
}
|
||
.container {
|
||
max-width: 1200px;
|
||
margin: 0 auto;
|
||
background: white;
|
||
border-radius: 16px;
|
||
box-shadow: 0 20px 60px rgba(0,0,0,0.1);
|
||
overflow: hidden;
|
||
}
|
||
.header {
|
||
background: linear-gradient(135deg, var(--primary) 0%, var(--secondary) 100%);
|
||
color: white;
|
||
padding: 40px;
|
||
text-align: center;
|
||
}
|
||
.header h1 { font-size: 2.1em; margin-bottom: 10px; }
|
||
.header .subtitle { font-size: 1.1em; opacity: 0.95; }
|
||
.meta-bar {
|
||
display: flex;
|
||
justify-content: space-between;
|
||
background: var(--bg-light);
|
||
padding: 15px 30px;
|
||
border-bottom: 2px solid var(--border);
|
||
flex-wrap: wrap;
|
||
gap: 15px;
|
||
}
|
||
.meta-item { display: flex; align-items: center; gap: 8px; font-size: 0.9em; color: var(--text-muted); }
|
||
.meta-item strong { color: var(--primary); }
|
||
.content { padding: 40px; }
|
||
.section {
|
||
margin-bottom: 35px;
|
||
padding: 25px;
|
||
background: var(--bg-light);
|
||
border-radius: 12px;
|
||
border-right: 5px solid var(--secondary);
|
||
}
|
||
.section h2 {
|
||
color: var(--primary);
|
||
font-size: 1.5em;
|
||
margin-bottom: 20px;
|
||
padding-bottom: 10px;
|
||
border-bottom: 2px solid var(--border);
|
||
}
|
||
.section h3 { color: var(--secondary); font-size: 1.2em; margin: 20px 0 12px; }
|
||
pre {
|
||
background: var(--bg-code);
|
||
color: #e2e8f0;
|
||
padding: 18px;
|
||
border-radius: 8px;
|
||
overflow-x: auto;
|
||
direction: ltr;
|
||
text-align: left;
|
||
font-family: 'Consolas', monospace;
|
||
font-size: 0.83em;
|
||
margin: 15px 0;
|
||
border-right: 4px solid var(--accent);
|
||
}
|
||
code {
|
||
background: #fef3c7;
|
||
color: #92400e;
|
||
padding: 2px 8px;
|
||
border-radius: 4px;
|
||
font-family: 'Consolas', monospace;
|
||
font-size: 0.9em;
|
||
direction: ltr;
|
||
display: inline-block;
|
||
}
|
||
table {
|
||
width: 100%;
|
||
border-collapse: collapse;
|
||
margin: 15px 0;
|
||
background: white;
|
||
border-radius: 8px;
|
||
overflow: hidden;
|
||
}
|
||
th { background: var(--primary); color: white; padding: 12px; text-align: right; }
|
||
td { padding: 12px; border-bottom: 1px solid var(--border); }
|
||
.alert { padding: 15px 20px; border-radius: 8px; margin: 15px 0; border-right: 4px solid; }
|
||
.alert-info { background: #dbeafe; border-color: var(--secondary); color: #1e40af; }
|
||
.alert-success { background: #d1fae5; border-color: var(--success); color: #065f46; }
|
||
.alert-warning { background: #fef3c7; border-color: var(--accent); color: #92400e; }
|
||
.footer { background: var(--primary); color: white; padding: 25px; text-align: center; }
|
||
.toc { background: white; padding: 20px; border-radius: 10px; margin-bottom: 25px; border: 2px solid var(--border); }
|
||
.toc h3 { color: var(--primary); margin-bottom: 15px; }
|
||
.toc ol { padding-right: 25px; }
|
||
.toc li { padding: 6px 0; }
|
||
.toc a { color: var(--secondary); text-decoration: none; }
|
||
.file-change { background: #f0f9ff; border-right: 4px solid var(--secondary); padding: 15px; margin: 10px 0; border-radius: 8px; }
|
||
.file-change .path { font-family: 'Consolas', monospace; color: var(--primary); font-weight: bold; direction: ltr; display: inline-block; }
|
||
.badge-new { background: var(--success); color: white; padding: 2px 8px; border-radius: 4px; font-size: 0.75em; margin-right: 8px; }
|
||
.badge-modify { background: var(--warning); color: white; padding: 2px 8px; border-radius: 4px; font-size: 0.75em; margin-right: 8px; }
|
||
.arch-grid { display: grid; grid-template-columns: repeat(auto-fit, minmax(280px, 1fr)); gap: 20px; margin: 20px 0; }
|
||
.arch-card { background: white; padding: 20px; border-radius: 10px; box-shadow: 0 4px 12px rgba(0,0,0,0.08); border-top: 4px solid var(--secondary); }
|
||
.arch-card h4 { color: var(--primary); margin-bottom: 12px; }
|
||
.arch-card ul { list-style: none; padding-right: 0; }
|
||
.arch-card li { padding: 6px 0; padding-right: 20px; position: relative; }
|
||
.arch-card li::before { content: '▸'; position: absolute; right: 0; color: var(--accent); font-weight: bold; }
|
||
</style>
|
||
</head>
|
||
<body>
|
||
<div class="container">
|
||
|
||
<div class="header">
|
||
<h1>👁️ استخراج محتوای جهانی با Qwen3.5-9B Vision</h1>
|
||
<div class="subtitle">یکپارچهسازی قابلیت OCR هوشمند در معماری XFileContentExtractor</div>
|
||
</div>
|
||
|
||
<div class="meta-bar">
|
||
<div class="meta-item">👨💻 <strong>توسعهدهنده:</strong> هادی خزاعی اصل</div>
|
||
<div class="meta-item">🏢 <strong>شرکت:</strong> فن آوران ساحر علم</div>
|
||
<div class="meta-item">📅 <strong>تاریخ:</strong> شنبه ۱۲ مهر ۱۴۰۵</div>
|
||
<div class="meta-item">📦 <strong>پروژه:</strong> xAiApi</div>
|
||
</div>
|
||
|
||
<div class="content">
|
||
|
||
<div class="toc">
|
||
<h3>📑 فهرست مطالب</h3>
|
||
<ol>
|
||
<li><a href="#capabilities">تحلیل قابلیتهای مدل Qwen3.5-9B GGUF</a></li>
|
||
<li><a href="#strategy">استراتژی استخراج جهانی (Universal Extraction)</a></li>
|
||
<li><a href="#step1">گام ۱: پیکربندی مدل در Ollama</a></li>
|
||
<li><a href="#step2">گام ۲: ایجاد XQwenVisionContentExtractor</a></li>
|
||
<li><a href="#step3">گام ۳: بهروزرسانی XPdfFileContentExtractor</a></li>
|
||
<li><a href="#step4">گام ۴: ثبت سرویسها در DI</a></li>
|
||
<li><a href="#step5">گام ۵: مهندسی پرامپت برای OCR دقیق</a></li>
|
||
<li><a href="#step6">گام ۶: ملاحظات عملکردی و Production</a></li>
|
||
</ol>
|
||
</div>
|
||
|
||
<!-- Section 1: Capabilities -->
|
||
<div class="section" id="capabilities">
|
||
<h2>🧠 گام ۱: تحلیل قابلیتهای مدل Qwen3.5-9B GGUF</h2>
|
||
<p>مدل <code>Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP-GGUF</code> یک مدل زبانی بزرگ کوانتیزه شده (GGUF) با ویژگیهای منحصر به فرد است:</p>
|
||
|
||
<div class="arch-grid">
|
||
<div class="arch-card">
|
||
<h4>👁️ قابلیت Vision / OCR</h4>
|
||
<ul>
|
||
<li>توانایی خواندن متن از تصاویر (PNG, JPG, WEBP)</li>
|
||
<li>تشخیص جداول و ساختارهای بصری در اسناد اسکنشده</li>
|
||
<li>حذف نیاز به کتابخانههای سنتی OCR مانند Tesseract</li>
|
||
</ul>
|
||
</div>
|
||
<div class="arch-card">
|
||
<h4>🚫 Uncensored & Heretic</h4>
|
||
<ul>
|
||
<li>بدون فیلترهای اخلاقی سختگیرانه</li>
|
||
<li>استخراج دقیق محتوا بدون "مoralizing" یا رد درخواست برای اسناد حساس</li>
|
||
<li>ایدهآل برای پردازش اسناد حقوقی، پزشکی یا فنی خام</li>
|
||
</ul>
|
||
</div>
|
||
<div class="arch-card">
|
||
<h4>⚡ IMATRIX & MTP</h4>
|
||
<ul>
|
||
<li>بهینهسازی شده برای سرعت استنتاج (Inference) بالاتر</li>
|
||
<li>پشتیبانی از Context Window گسترده (معمولاً 8K تا 32K توکن)</li>
|
||
<li>مناسب برای پردازش اسناد چند صفحهای</li>
|
||
</ul>
|
||
</div>
|
||
</div>
|
||
</div>
|
||
|
||
<!-- Section 2: Strategy -->
|
||
<div class="section" id="strategy">
|
||
<h2>🎯 گام ۲: استراتژی استخراج جهانی (Universal Extraction)</h2>
|
||
<p>به جای استفاده از Extractor های جداگانه و پیچیده برای هر فرمت، یک <strong>مسیر هوشمند ترکیبی</strong> طراحی میکنیم:</p>
|
||
|
||
<table>
|
||
<tr>
|
||
<th>نوع فایل</th>
|
||
<th>استراتژی استخراج</th>
|
||
<th>مسیر پردازش</th>
|
||
</tr>
|
||
<tr>
|
||
<td>📄 PDF متنی (Native)</td>
|
||
<td>استخراج متن سریع با PdfPig</td>
|
||
<td>XPdfFileContentExtractor (متن) → بازگشت سریع</td>
|
||
</tr>
|
||
<tr>
|
||
<td>📄 PDF اسکنشده / تصویری</td>
|
||
<td>تبدیل صفحات به تصویر + OCR با Qwen</td>
|
||
<td>XPdfFileContentExtractor (تصویر) → XQwenVisionContentExtractor</td>
|
||
</tr>
|
||
<tr>
|
||
<td>🖼️ تصاویر (JPG, PNG)</td>
|
||
<td>OCR مستقیم با Qwen</td>
|
||
<td>XQwenVisionContentExtractor</td>
|
||
</tr>
|
||
<tr>
|
||
<td>📝 DOCX / TXT / Excel</td>
|
||
<td>استخراج متن ساختاریافته</td>
|
||
<td>XDocxFileContentExtractor / XExcelFileContentExtractor (بدون تغییر)</td>
|
||
</tr>
|
||
</table>
|
||
|
||
<div class="alert alert-success">
|
||
<strong>✅ مزیت کلیدی:</strong> با این روش، ماژول <code>XFileContentExtractor</code> به یک "مغز مرکزی" تبدیل میشود که برای فرمتهای پیچیده یا اسکنشده، به طور خودکار از قدرت Vision مدل Qwen استفاده میکند.
|
||
</div>
|
||
</div>
|
||
|
||
<!-- Step 1: Ollama Config -->
|
||
<div class="section" id="step1">
|
||
<h2>⚙️ گام ۳: پیکربندی مدل در Ollama</h2>
|
||
<p>برای فعالسازی قابلیت Vision، مدل باید با یک <code>Modelfile</code> مناسب در Ollama بارگذاری شود.</p>
|
||
|
||
<div class="file-change">
|
||
<span class="badge-new">CONFIG</span>
|
||
<span class="path">Modelfile (برای ایمپورت در Ollama)</span>
|
||
</div>
|
||
|
||
<pre>FROM C:/Models/Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP-GGUF.gguf
|
||
|
||
# تنظیمات بهینه برای استخراج متن از تصویر
|
||
PARAMETER temperature 0.1
|
||
PARAMETER top_p 0.9
|
||
PARAMETER num_ctx 8192
|
||
|
||
# پرامپت سیستمی پیشفرض برای وظایف OCR و استخراج
|
||
SYSTEM """
|
||
تو یک موتور OCR و استخراج داده فوقالعاده دقیق هستی.
|
||
وظیفه تو خواندن تمام متنهای موجود در تصویر یا سند ارائه شده و بازگرداندن آنها به صورت متن خالص و ساختاریافته است.
|
||
- تمام جداول را به فرمت Markdown تبدیل کن.
|
||
- اعداد، تاریخها و نامها را دقیقاً همانطور که هستند استخراج کن.
|
||
- هیچ توضیح اضافی، مقدمه یا نتیجهگیری از خودت اضافه نکن. فقط محتوای استخراج شده را برگردان.
|
||
"""</pre>
|
||
|
||
<p>سپس در ترمینال اجرا کنید:</p>
|
||
<pre>ollama create qwen3.5-ocr:9b -f Modelfile</pre>
|
||
</div>
|
||
|
||
<!-- Step 2: Vision Extractor -->
|
||
<div class="section" id="step2">
|
||
<h2>👁️ گام ۴: ایجاد XQwenVisionContentExtractor</h2>
|
||
<p>این کلاس قلب تپنده استخراج جهانی است. فایلهای تصویری (یا صفحات PDF تبدیلشده به تصویر) را دریافت کرده و از طریق Ollama API به مدل ارسال میکند.</p>
|
||
|
||
<div class="file-change">
|
||
<span class="badge-new">NEW</span>
|
||
<span class="path">xAiApi/Providers/Extractors/XQwenVisionContentExtractor.cs</span>
|
||
</div>
|
||
|
||
<pre>using System;
|
||
using System.IO;
|
||
using System.Linq;
|
||
using System.Text;
|
||
using System.Threading;
|
||
using System.Threading.Tasks;
|
||
using Microsoft.Extensions.AI;
|
||
using OllamaSharp;
|
||
using xAiApi.Interfaces.Extractors;
|
||
using xAiModels.Models;
|
||
|
||
namespace xAiApi.Providers.Extractors
|
||
{
|
||
/// <summary>
|
||
/// استخراج محتوا از فایلهای تصویری با استفاده از قابلیت Vision مدل Qwen ...
|
||
/// </summary>
|
||
public class XQwenVisionContentExtractor : IXFileContentExtractor
|
||
{
|
||
private static readonly string[] SupportedMimeTypes =
|
||
[
|
||
"image/png", "image/jpeg", "image/jpg", "image/webp", "image/bmp",
|
||
"application/pdf" // برای هندل کردن PDF های اسکن شده در سطح Vision
|
||
];
|
||
|
||
private readonly string _ollamaUrl;
|
||
private readonly string _modelName;
|
||
private readonly string _extractionPrompt;
|
||
|
||
public XQwenVisionContentExtractor(
|
||
string ollamaUrl = "http://localhost:11434",
|
||
string modelName = "qwen3.5-ocr:9b",
|
||
string extractionPrompt = "تمام متن موجود در این تصویر را با دقت بالا استخراج کن و به صورت متن خالص برگردان.")
|
||
{
|
||
_ollamaUrl = ollamaUrl;
|
||
_modelName = modelName;
|
||
_extractionPrompt = extractionPrompt;
|
||
}
|
||
|
||
public bool CanExtract(string mimeType)
|
||
{
|
||
return SupportedMimeTypes.Contains(mimeType?.ToLowerInvariant() ?? string.Empty);
|
||
}
|
||
|
||
public async Task<string> ExtractAsync(
|
||
Stream fileStream,
|
||
string mimeType,
|
||
CancellationToken cancellationToken = default)
|
||
{
|
||
var result = await ExtractRichAsync(fileStream, "file", mimeType, cancellationToken);
|
||
return result.Text;
|
||
}
|
||
|
||
public async Task<XFileExtractionResult> ExtractRichAsync(
|
||
Stream fileStream,
|
||
string fileName,
|
||
string mimeType,
|
||
CancellationToken cancellationToken = default)
|
||
{
|
||
var result = new XFileExtractionResult
|
||
{
|
||
FileName = fileName,
|
||
MimeType = mimeType
|
||
};
|
||
|
||
try
|
||
{
|
||
// ۱. خواندن استریم به صورت بایت (برای ارسال به مدل Vision)
|
||
using var memoryStream = new MemoryStream();
|
||
await fileStream.CopyToAsync(memoryStream, cancellationToken);
|
||
var imageBytes = memoryStream.ToArray();
|
||
|
||
// ۲. تنظیم کلاینت Ollama
|
||
var ollamaClient = new OllamaApiClient(new Uri(_ollamaUrl), _modelName);
|
||
|
||
// ۳. ساخت پیام چندوجهی (Text + Image)
|
||
var messages = new[]
|
||
{
|
||
new ChatMessage(ChatRole.User, new[]
|
||
{
|
||
new TextContent(_extractionPrompt),
|
||
new DataContent(imageBytes, mimeType) // ارسال تصویر به مدل
|
||
})
|
||
};
|
||
|
||
// ۴. فراخوانی مدل برای استخراج متن
|
||
var response = await ollamaClient.GetResponseAsync(messages, cancellationToken: cancellationToken);
|
||
|
||
if (!string.IsNullOrWhiteSpace(response.Text))
|
||
{
|
||
result.Text = response.Text.Trim();
|
||
result.IsValid = true;
|
||
}
|
||
else
|
||
{
|
||
result.ErrorMessage = "مدل پاسخی برای استخراج تولید نکرد.";
|
||
}
|
||
}
|
||
catch (Exception ex)
|
||
{
|
||
result.ErrorMessage = $"خطا در استخراج Vision: {ex.Message}";
|
||
}
|
||
|
||
return result;
|
||
}
|
||
}
|
||
}</pre>
|
||
|
||
<div class="alert alert-info">
|
||
<strong>💡 نکته فنی:</strong> کلاس <code>DataContent</code> از <code>Microsoft.Extensions.AI</code> به طور خودکار بایتهای تصویر را به فرمت Base64 مناسب برای API مدلهای Vision تبدیل میکند.
|
||
</div>
|
||
</div>
|
||
|
||
<!-- Step 3: PDF Update -->
|
||
<div class="section" id="step3">
|
||
<h2>📄 گام ۵: بهروزرسانی XPdfFileContentExtractor</h2>
|
||
<p>اکنون منطق Fallback را تغییر میدهیم. به جای اینکه فقط تصویر را ذخیره کنیم، آن تصویر را به <code>XQwenVisionContentExtractor</code> میفرستیم تا متن را استخراج کند.</p>
|
||
|
||
<div class="file-change">
|
||
<span class="badge-modify">MODIFY</span>
|
||
<span class="path">xAiApi/Providers/Extractors/XPdfFileContentExtractor.cs</span>
|
||
</div>
|
||
|
||
<pre>using System;
|
||
using System.IO;
|
||
using System.Linq;
|
||
using System.Text;
|
||
using System.Threading;
|
||
using System.Threading.Tasks;
|
||
using UglyToad.PdfPig;
|
||
using xAiApi.Interfaces.Extractors;
|
||
using xAiModels.Models;
|
||
|
||
namespace xAiApi.Providers.Extractors
|
||
{
|
||
public class XPdfFileContentExtractor : IXFileContentExtractor
|
||
{
|
||
private const int MinTextLengthThreshold = 100;
|
||
private const int MaxPagesForImageFallback = 5; // محدود کردن برای جلوگیری از مصرف بیش از حد GPU
|
||
|
||
private readonly IXFileContentExtractor _visionExtractor;
|
||
|
||
// تزریق Vision Extractor از طریق Constructor
|
||
public XPdfFileContentExtractor(IXFileContentExtractor visionExtractor)
|
||
{
|
||
_visionExtractor = visionExtractor;
|
||
}
|
||
|
||
public bool CanExtract(string mimeType) => mimeType?.ToLowerInvariant() == "application/pdf";
|
||
|
||
public async Task<string> ExtractAsync(Stream fileStream, string mimeType, CancellationToken cancellationToken = default)
|
||
{
|
||
var result = await ExtractRichAsync(fileStream, "document.pdf", mimeType, cancellationToken);
|
||
return result.Text;
|
||
}
|
||
|
||
public async Task<XFileExtractionResult> ExtractRichAsync(
|
||
Stream fileStream,
|
||
string fileName,
|
||
string mimeType,
|
||
CancellationToken cancellationToken = default)
|
||
{
|
||
var result = new XFileExtractionResult { FileName = fileName, MimeType = mimeType };
|
||
|
||
using var memoryStream = new MemoryStream();
|
||
await fileStream.CopyToAsync(memoryStream, cancellationToken);
|
||
|
||
// مرحله ۱: تلاش برای استخراج متن معمولی
|
||
memoryStream.Position = 0;
|
||
var extractedText = await ExtractTextAsync(memoryStream, cancellationToken);
|
||
|
||
if (IsTextSufficient(extractedText))
|
||
{
|
||
result.Text = extractedText;
|
||
result.IsValid = true;
|
||
return result;
|
||
}
|
||
|
||
// مرحله ۲: Fallback به Vision OCR
|
||
memoryStream.Position = 0;
|
||
result = await FallbackToVisionOcrAsync(memoryStream, fileName, cancellationToken);
|
||
result.UsedImageFallback = true;
|
||
|
||
return result;
|
||
}
|
||
|
||
private async Task<string> ExtractTextAsync(Stream pdfStream, CancellationToken cancellationToken)
|
||
{
|
||
return await Task.Run(() =>
|
||
{
|
||
var sb = new StringBuilder();
|
||
using var document = PdfDocument.Open(pdfStream);
|
||
foreach (var page in document.GetPages())
|
||
{
|
||
sb.AppendLine(page.Text?.Trim());
|
||
}
|
||
return sb.ToString();
|
||
}, cancellationToken);
|
||
}
|
||
|
||
private bool IsTextSufficient(string text)
|
||
{
|
||
if (string.IsNullOrWhiteSpace(text)) return false;
|
||
var cleanText = new string(text.Where(c => !char.IsWhiteSpace(c)).ToArray());
|
||
return cleanText.Length >= MinTextLengthThreshold;
|
||
}
|
||
|
||
private async Task<XFileExtractionResult> FallbackToVisionOcrAsync(
|
||
Stream pdfStream,
|
||
string fileName,
|
||
CancellationToken cancellationToken)
|
||
{
|
||
var result = new XFileExtractionResult { FileName = fileName, MimeType = "application/pdf", UsedImageFallback = true };
|
||
var sb = new StringBuilder();
|
||
|
||
await Task.Run(() =>
|
||
{
|
||
using var document = PdfDocument.Open(pdfStream);
|
||
var pageCount = Math.Min(document.NumberOfPages, MaxPagesForImageFallback);
|
||
|
||
for (int i = 0; i < pageCount; i++)
|
||
{
|
||
// تبدیل صفحه به تصویر (با استفاده از PdfPig یا Pdfium)
|
||
// نکته: برای سادگی، فرض میکنیم متدی داریم که صفحه را به Stream تصویر تبدیل میکند
|
||
// در پیادهسازی واقعی از PdfiumViewer.RenderPage استفاده کنید (همانند کد قبلی)
|
||
using var pageImageStream = RenderPageToImageStream(document, i);
|
||
|
||
// ارسال تصویر به Qwen Vision Extractor
|
||
var ocrResult = _visionExtractor.ExtractRichAsync(
|
||
pageImageStream,
|
||
$"{fileName}_page_{i+1}",
|
||
"image/png",
|
||
cancellationToken).GetAwaiter().GetResult();
|
||
|
||
if (ocrResult.IsValid)
|
||
{
|
||
sb.AppendLine($"--- Page {i + 1} ---");
|
||
sb.AppendLine(ocrResult.Text);
|
||
}
|
||
}
|
||
}, cancellationToken);
|
||
|
||
result.Text = sb.ToString();
|
||
result.IsValid = !string.IsNullOrWhiteSpace(result.Text);
|
||
return result;
|
||
}
|
||
|
||
// Placeholder: باید با منطق PdfiumViewer که قبلاً نوشتید جایگزین شود
|
||
private Stream RenderPageToImageStream(dynamic document, int pageIndex)
|
||
{
|
||
// پیادهسازی RenderPage به MemoryStream از PdfiumViewer
|
||
throw new NotImplementedException("از منطق PdfiumViewer.RenderPage استفاده کنید");
|
||
}
|
||
}
|
||
}</pre>
|
||
</div>
|
||
|
||
<!-- Step 4: DI Registration -->
|
||
<div class="section" id="step4">
|
||
<h2>🔌 گام ۶: ثبت سرویسها در Dependency Injection</h2>
|
||
<p>اکنون باید زنجیره Extractor ها را در <code>Startup.cs</code> به درستی متصل کنیم.</p>
|
||
|
||
<div class="file-change">
|
||
<span class="badge-modify">MODIFY</span>
|
||
<span class="path">xAiApi/Startup.cs (متد ConfigureServices)</span>
|
||
</div>
|
||
|
||
<pre>public void ConfigureServices(IServiceCollection services)
|
||
{
|
||
// ... (ثبتهای قبلی)
|
||
|
||
// ۱. ثبت Vision Extractor به صورت Singleton (چون State-less است)
|
||
services.AddSingleton<IXFileContentExtractor>(sp =>
|
||
new XQwenVisionContentExtractor(
|
||
ollamaUrl: "http://localhost:11434",
|
||
modelName: "qwen3.5-ocr:9b"
|
||
));
|
||
|
||
// ۲. ثبت PDF Extractor و تزریق Vision Extractor به آن
|
||
services.AddSingleton<IXFileContentExtractor>(sp =>
|
||
{
|
||
var visionExtractor = sp.GetRequiredService<IXFileContentExtractor>();
|
||
// نکته: برای جلوگیری از تداخل در Resolve، بهتر است Vision Extractor را با نام/interface خاص ثبت کنید
|
||
// یا مستقیماً اینstantiate کنید:
|
||
var specificVisionExtractor = new XQwenVisionContentExtractor();
|
||
return new XPdfFileContentExtractor(specificVisionExtractor);
|
||
});
|
||
|
||
// ۳. ثبت سایر Extractor ها
|
||
services.AddSingleton<IXFileContentExtractor, XPlainTextFileContentExtractor>();
|
||
services.AddSingleton<IXFileContentExtractor, XDocxFileContentExtractor>();
|
||
services.AddSingleton<IXFileContentExtractor, XExcelFileContentExtractor>();
|
||
|
||
// ۴. ثبت Composite Extractor (این کلاس به طور خودکار همه IXFileContentExtractor های ثبت شده را جمعآوری میکند)
|
||
services.AddSingleton<XFileContentExtractor>();
|
||
|
||
// اطمینان از اینکه سرویس اصلی از نوع Composite است
|
||
services.AddSingleton<IXFileContentExtractor>(sp => sp.GetRequiredService<XFileContentExtractor>());
|
||
|
||
// ... (بقیه کدها)
|
||
}</pre>
|
||
</div>
|
||
|
||
<!-- Step 5: Prompt Engineering -->
|
||
<div class="section" id="step5">
|
||
<h2>📝 گام ۷: مهندسی پرامپت برای OCR دقیق</h2>
|
||
<p>کیفیت استخراج مستقیماً به پرامپت ارسال شده به مدل Qwen بستگی دارد. این پرامپتها را در <code>appsettings.json</code> یا کد تعریف کنید:</p>
|
||
|
||
<div class="arch-grid">
|
||
<div class="arch-card">
|
||
<h4>🎯 پرامپت استخراج عمومی (General OCR)</h4>
|
||
<p style="font-size: 0.9em; color: var(--text-muted);">
|
||
"تمام متن موجود در این تصویر را با دقت کاراکتر به کاراکتر استخراج کن. ساختار پاراگرافها را حفظ کن. اگر جدولی وجود دارد، آن را به فرمت Markdown تبدیل کن. هیچ توضیح اضافی نده."
|
||
</p>
|
||
</div>
|
||
<div class="arch-card">
|
||
<h4>📊 پرامپت استخراج داده ساختاریافته (Structured)</h4>
|
||
<p style="font-size: 0.9em; color: var(--text-muted);">
|
||
"این تصویر یک فاکتور/سند است. فقط موارد زیر را استخراج و به صورت JSON برگردان: {\"issuer\": \"\", \"date\": \"\", \"total_amount\": \"\", \"items\": []}. اگر موردی یافت نشد، null بگذار."
|
||
</p>
|
||
</div>
|
||
</div>
|
||
</div>
|
||
|
||
<!-- Step 6: Production Considerations -->
|
||
<div class="section" id="step6">
|
||
<h2>⚠️ گام ۸: ملاحظات عملکردی و Production</h2>
|
||
|
||
<table>
|
||
<tr>
|
||
<th>چالش</th>
|
||
<th>راهحل پیشنهادی در معماری</th>
|
||
</tr>
|
||
<tr>
|
||
<td>محدودیت Context Window</td>
|
||
<td>در <code>XPdfFileContentExtractor</code>، <code>MaxPagesForImageFallback</code> را روی ۵ تا ۱۰ تنظیم کنید تا مدل Overload نشود.</td>
|
||
</tr>
|
||
<tr>
|
||
<td>زمان پاسخدهی (Latency)</td>
|
||
<td>پردازش Vision زمانبر است. در <code>AskAsync</code> از Timeout مناسب (مثلاً ۶۰ ثانیه) در HttpClient استفاده کنید.</td>
|
||
</tr>
|
||
<tr>
|
||
<td>مصرف حافظه GPU</td>
|
||
<td>مدل 9B حدود ۶-۸ گیگابایت VRAM نیاز دارد. اطمینان حاصل کنید سرور Ollama منابع کافی دارد. از پردازش همزمان بیش از ۲ فایل بزرگ خودداری کنید.</td>
|
||
</tr>
|
||
<tr>
|
||
<td>کیفیت OCR زبان فارسی</td>
|
||
<td>مدلهای Qwen در فارسی خوب عمل میکنند، اما برای اسناد بسیار قدیمی یا با کیفیت پایین، ممکن است نیاز به پیشپردازش تصویر (افزایش کنتراست) باشد.</td>
|
||
</tr>
|
||
</table>
|
||
|
||
<div class="alert alert-warning">
|
||
<strong>⚠️ نکته حیاتی درباره XFileContentExtractor فعلی:</strong><br>
|
||
کد فعلی <code>XFileContentExtractor</code> از Reflection برای یافتن Extractor ها استفاده میکند. برای اطمینان از عملکرد صحیح، اطمینان حاصل کنید که <code>XQwenVisionContentExtractor</code> و <code>XPdfFileContentExtractor</code> در همان Assembly (پروژه xAiApi) کامپایل شدهاند تا توسط Reflection شناسایی شوند.
|
||
</div>
|
||
</div>
|
||
|
||
</div>
|
||
|
||
<div class="footer">
|
||
<p><strong>👨💻 توسعهدهنده:</strong> هادی خزاعی اصل</p>
|
||
<p><strong>🏢 شرکت:</strong> فن آوران ساحر علم</p>
|
||
<p><strong>📅 تاریخ:</strong> شنبه ۱۲ مهر ۱۴۰۵</p>
|
||
<p style="margin-top: 15px; opacity: 0.8; font-size: 0.9em;">
|
||
👁️ مستند فنی استخراج جهانی با Qwen3.5-9B Vision - تمامی حقوق محفوظ است
|
||
</p>
|
||
</div>
|
||
|
||
</div>
|
||
</body>
|
||
</html> |