
Mario Magdy
Cairo, Egypt
Mario Magdy
Data Scientist and AI Engineer
Category : Artificial intelligence (AI)
I’m an AI Engineer and LLM Systems Builder focused on designing reliable, production-grade AI pipelines that bridge natural language, data, and decision-making.
I’m a Top Rated Upwork freelancer (100% Job Success, 1,400+ hours worked), where I’ve delivered AI and data systems for U.S.-based clients across real estate, finance, and analytics domains. My work spans LLM-driven document analysis, automated data enrichment pipelines, and backend systems built with Python, FastAPI, and Flask.
At GenLift, I design and ship end-to-end LLM workflows that convert natural language basketball queries into structured analytical outputs. I built a formal LLM evaluation harness that measures accuracy and stability across paraphrases, implemented regression testing for prompt robustness, and reduced hallucination rates through structured validation pipelines.
For real estate investors, I engineered multi-stage automation systems that:
• Scrape probate court records
• Use Generative AI to extract structured insights from PDFs
• Perform identity resolution and contact enrichment
• Deliver actionable lead datasets
In finance, I developed LLM-based stock analysis pipelines that normalize press releases and generate structured investment summaries with controlled outputs and model selection safeguards.
Previously, I co-founded DeepScope, an AI-powered smart telescope startup that won 1st place at the Egyptian Space Summit for technical innovation.
My core expertise includes:
• LLM pipelines & evaluation frameworks
• AI system reliability & testing
• FastAPI / Flask backend engineering
• Data automation & enrichment systems
• Prompt engineering with measurable validation
• Production-grade Python systems
I focus on building AI systems that are not just impressive demos — but measurable, testable, and deployable in real environments.
I’m a Top Rated Upwork freelancer (100% Job Success, 1,400+ hours worked), where I’ve delivered AI and data systems for U.S.-based clients across real estate, finance, and analytics domains. My work spans LLM-driven document analysis, automated data enrichment pipelines, and backend systems built with Python, FastAPI, and Flask.
At GenLift, I design and ship end-to-end LLM workflows that convert natural language basketball queries into structured analytical outputs. I built a formal LLM evaluation harness that measures accuracy and stability across paraphrases, implemented regression testing for prompt robustness, and reduced hallucination rates through structured validation pipelines.
For real estate investors, I engineered multi-stage automation systems that:
• Scrape probate court records
• Use Generative AI to extract structured insights from PDFs
• Perform identity resolution and contact enrichment
• Deliver actionable lead datasets
In finance, I developed LLM-based stock analysis pipelines that normalize press releases and generate structured investment summaries with controlled outputs and model selection safeguards.
Previously, I co-founded DeepScope, an AI-powered smart telescope startup that won 1st place at the Egyptian Space Summit for technical innovation.
My core expertise includes:
• LLM pipelines & evaluation frameworks
• AI system reliability & testing
• FastAPI / Flask backend engineering
• Data automation & enrichment systems
• Prompt engineering with measurable validation
• Production-grade Python systems
I focus on building AI systems that are not just impressive demos — but measurable, testable, and deployable in real environments.
Working hours
- Monday:08h00 To 18h00
- Tuesday:08h00 To 18h00
- Wednesday:08h00 To 18h00
- Thursday:08h00 To 18h00
- Friday:08h00 To 18h00
- Saturday:Not available
- Sunday:Not available
• Delivered an end-to-end workflow that turns natural-language requests into extracted parameters,
executes the request through an LLM layer (Flask/Pandas), and returns explainable outputs and video
context, working closely with a cross-functional team. Validated the migration by running two parallel
branches (legacy AssistJS+Flask and the new FastAPI branch) side by side.
• Built a team-owned LLM evaluation harness (v0.1 to v0.2) using Group Questions and voting-based
consensus to approximate ground truth when labels were missing, and formalized two metrics (Accuracy
and Stability across paraphrases) to guide model selection. Reduced hallucinations by 65%.
• Implemented prompt-robustness regression tests by generating paraphrase groups per intent and
enforcing that different wording leads to the same intent and the same retrieval/query logic.
• Scaled stress testing with a Griller flow that generates 9 rephrase variations per query, runs them in
parallel, and stores raw and filtered outputs for review, achieving 10x paraphrase coverage and persistent
evaluation data for QA and regression tracking.
executes the request through an LLM layer (Flask/Pandas), and returns explainable outputs and video
context, working closely with a cross-functional team. Validated the migration by running two parallel
branches (legacy AssistJS+Flask and the new FastAPI branch) side by side.
• Built a team-owned LLM evaluation harness (v0.1 to v0.2) using Group Questions and voting-based
consensus to approximate ground truth when labels were missing, and formalized two metrics (Accuracy
and Stability across paraphrases) to guide model selection. Reduced hallucinations by 65%.
• Implemented prompt-robustness regression tests by generating paraphrase groups per intent and
enforcing that different wording leads to the same intent and the same retrieval/query logic.
• Scaled stress testing with a Griller flow that generates 9 rephrase variations per query, runs them in
parallel, and stores raw and filtered outputs for review, achieving 10x paraphrase coverage and persistent
evaluation data for QA and regression tracking.
• Started with a market-screening phase to choose where to invest first, using publicly available signals
across cities, counties, and metros to prioritize the initial target region before lead generation.
• Designed a two-source enrichment workflow to convert public records into actionable leads by combining
court and probate case discovery with FastPeopleSearch identity resolution (names, addresses, relatives),
producing one consolidated lead record per case. Result: an auditable pipeline from court records to
people search to a merged lead, ready for CSV and Excel export; volume processed and match-rate
improvement:
• Imported structured leads into Zoho CRM so the sales and outreach team could operate from a single
pipeline.
• Established an operating process in Zoho CRM to continuously ingest new leads, track statuses, and
maintain clean data for segmentation, outreach, and ongoing analysis.
across cities, counties, and metros to prioritize the initial target region before lead generation.
• Designed a two-source enrichment workflow to convert public records into actionable leads by combining
court and probate case discovery with FastPeopleSearch identity resolution (names, addresses, relatives),
producing one consolidated lead record per case. Result: an auditable pipeline from court records to
people search to a merged lead, ready for CSV and Excel export; volume processed and match-rate
improvement:
• Imported structured leads into Zoho CRM so the sales and outreach team could operate from a single
pipeline.
• Established an operating process in Zoho CRM to continuously ingest new leads, track statuses, and
maintain clean data for segmentation, outreach, and ongoing analysis.
Engineering Analysis, Astronomy, Space Technology, Artificial Intelligence, Programming
- 🇬🇧 English
- 🇲🇦 Arabic
Please sign in as a customer to give your feedback




