I work where AI safety policy meets production reality: model outputs, reviewer judgment, escalation pressure, and sensitive-content risk. My job has usually been to find the part of the system that is quietly drifting and make it measurable enough to fix.
AI Trust & Safety and compliance professional with 6+ years across AI model evaluation, red-team prompt review, child safety, fraud, e-commerce risk, and large-scale moderation operations for TikTok, Meta programs, Highspring, and Handshake AI.
I have worked the full chain: frontline review, QA calibration, SME escalation, team management, vendor quality, LLM prompt evaluation, and policy alignment. That matters because the real problems rarely sit in one box. A model miss can be a policy problem. A policy miss can become a training problem. A training problem can become a queue-health problem by Friday afternoon.
My work is strongest in ambiguous, high-sensitivity environments: child safety, self-harm, sextortion, spam/scam, fraud, and AI outputs that need careful review rather than a fast label.
Languages: English, Arabic
Evaluate high-sensitivity AI prompts and model responses where policy, legal escalation, and human judgment all have to line up. The work is deliberately careful: the goal is not fast labeling, but consistent safety judgment under pressure.
Support AI safety and compliance evaluation for enterprise AI systems, with a focus on prompt red-teaming, model-output review, evaluator calibration, and root-cause analysis when guidance is being applied inconsistently.
Worked at the intersection of AI moderation, seller compliance, QA systems, vendor performance, and policy enforcement in a high-volume social commerce environment.
Led high-risk moderation operations supporting Meta platforms, balancing speed, accuracy, and policy consistency across sensitive content categories.
Served as escalation point and systems thinker across moderation operations, bridging frontline reviewers, QA, and leadership on complex policy decisions.
Focused on reviewer accuracy, evaluation consistency, and policy calibration in a fast-moving moderation environment where quality drift could quickly affect enforcement outcomes.
Worked directly in frontline review queues, where repeated exposure to abuse patterns, reviewer variation, and edge-case content built the foundation for later QA and systems-focused work.
This case study is based on the same kind of operating pattern I have seen across AI evaluation, moderation QA, and high-risk Trust & Safety work: the system can look healthy in aggregate while the dangerous misses collect in small, ambiguous categories. That is where I spend most of my attention.
Prompts and model responses involving minors, sexualized language, grooming indicators, self-harm, or coercion can fail quietly when the review process treats them like ordinary policy calls. The language may be indirect. The risk may sit in context. The correct action may require escalation discipline, not just a label.
Most safety policies are written to be precise, but production cases arrive messy. Reviewers and models can both miss risk when they over-index on explicit terms and underweight context, pattern, age signals, or coercive intent.
AI safety evaluation needs category-level controls, not just overall accuracy. For sensitive content, the work is to make the escalation threshold clear enough that careful reviewers and automated systems stop drifting in opposite directions.
Illustrative pattern based on operational review experience, not a published dataset.
QA finds the miss after the decision
Guidance flags risk before routing fails
For child safety and self-harm categories, earlier control design matters more than prettier dashboard numbers.
Model responses or evaluator decisions can appear policy-compliant at the surface while missing the actual enforcement intent. I saw this most clearly in LLM QA and compliance evaluation work, where a prompt can be answered cleanly but still be unsafe, evasive, or misaligned with the policy objective.
The failure is usually not that the policy is absent. The failure is translation. Policy language, prompt design, reviewer rubrics, and model behavior each carry their own assumptions. If nobody reconciles those assumptions, QA becomes late-stage cleanup.
A useful evaluation loop separates the error type before proposing a fix: bad prompt, bad model behavior, unclear policy, weak rubric, evaluator calibration, or escalation threshold. That is how you get from “accuracy is down” to a concrete repair.
Fraud, spam/scam, seller abuse, and coordinated manipulation rarely show up as one perfect signal. They show up as weak signals that become meaningful only when clustered: account history, language reuse, timing, seller behavior, payment risk, prior enforcement, and reviewer notes.
Queues and dashboards are often optimized for volume. Risk work needs prioritization. When systems treat every signal as equal, the important thing is easy to bury under the merely noisy thing.
Good Trust & Safety operations depend on evidence weighting. The job is not to collect every signal. The job is to decide which signals deserve action, review, escalation, or better instrumentation.
Every safety system has a threshold problem. Push too hard and you create false positives, appeals, reviewer friction, and business impact. Push too softly and harmful content remains active. The work is not pretending the tradeoff is avoidable. The work is knowing which categories cannot tolerate the same threshold.
Incorrect enforcement against non-violating content rises as rules get stricter.
Missed harmful content rises as rules get looser.
Friction increases when ambiguous content is pushed through stricter enforcement thresholds.
Exposure increases when looser enforcement allows more harmful content to remain active.
I tend to look for the control that is missing or too vague. In AI safety and moderation work, that usually means tracing a bad outcome backward until the failure is specific enough to repair.
That is the perspective I bring to AI evaluation, compliance QA, moderation quality, and Trust & Safety operations.
This case reflects work across LLM QA, red-team prompt review, and vendor quality audits where model output, evaluator reasoning, and policy intent all had to align under production constraints.
Evaluators could reach different decisions on the same prompt-response pair because the rubric did not clearly separate harmful content, allowed discussion, evasive model behavior, and escalation-worthy ambiguity.
Policies are written to define intent, but evaluation work needs operational examples, boundary cases, and repeatable reasoning. Without that, the same policy becomes several different policies in practice.
Decisions depend on reviewer interpretation
Reasoning is easier to audit and coach
QA is not just validation. In AI safety work, it is a control system for policy interpretation, evaluator behavior, and model feedback quality.
Based on moderation and compliance work across child safety, sextortion, spam/scam, fraud, and sensitive-content queues where the wrong routing decision can matter more than the initial label.
Reviewers often knew a case felt risky but did not always have a clean path for documenting why, escalating it, or distinguishing a policy violation from a legal or safety escalation trigger.
Escalation guidance is often written after the policy, when it should be part of the policy design. If reviewers have to invent the path during live work, the system is already taking unnecessary risk.
Escalation depends on individual judgment
Risk is routed with evidence and urgency
Strong escalation design is a safety mechanism. It protects users, reviewers, vendors, and the platform by making high-risk decisions consistent enough to defend.
Austin Community College
Scrum Alliance
PeopleCert
Google / Coursera
Amazon Web Services
CSSC
CSSC
Google / Coursera
SuccessCOACHING
University of California, Irvine / Coursera
DeepLearning.AI
Professional Development
Securiti
Evaluate prompts and model responses for safety, compliance, child-safety risk, self-harm risk, harmful output patterns, and policy alignment across current AI evaluation programs.
Managed Meta program review teams across high-risk queues, maintained SLA performance, scaled Arabic-language counterterrorism operations, and built decision structures that improved first-pass accuracy.
Built QA loops, evaluator coaching, knowledge bases, dashboards, and vendor review rhythms that turned recurring mistakes into clearer controls and measurable improvement.