About the job Remote | MCP & AI Connector Evaluation Specialist — $45–$185/hour
We are sharing a specialised part-time consulting opportunity for advanced LLM users with strong hands-on experience using Model Context Protocol tools, plugins, and connectors for complex personal workflows.
This role supports an AI research initiative focused on evaluating how effectively AI assistants complete personalised, multi-step tasks using connected tools such as Google Drive, Notion, travel platforms, and other plugins or connectors. Selected professionals will create realistic tasks, execute workflows while recording their screens, assess model performance, and develop detailed evaluation rubrics grounded in practical everyday use.
Key Responsibilities
Personal Workflow Design
- Create realistic prompts involving complex, high-context personal tasks
- Develop scenarios across travel, health research, dining, activity planning, home services, career search, and personal organisation
- Incorporate genuine preferences, constraints, trade-offs, and success criteria
- Design tasks requiring planning, judgment, context retention, and connected-tool usage
MCP, Plugin & Connector Testing
- Use MCP-enabled tools, plugins, and connectors to complete multi-step workflows
- Test integrations involving Google Drive, Notion, travel services, and similar platforms
- Evaluate whether AI systems select and use connected tools appropriately
- Identify failures involving permissions, context retrieval, sequencing, or action execution
Screen-Recorded Task Execution
- Complete assigned workflows while recording the screen
- Clearly demonstrate the actions, tools, and decisions involved in each task
- Document where the AI succeeds, overreaches, misses context, or produces impractical results
- Complete tasks within required turnaround windows
Model Evaluation & Rubric Development
- Judge whether outputs are personalised, realistic, useful, safe, and well-reasoned
- Write clear explanations of model strengths, weaknesses, and failure patterns
- Create detailed scoring rubrics for complex personal-assistant tasks
- Apply evaluation criteria consistently across model outputs
- Identify incomplete reasoning, unrealistic recommendations, and incorrect tool use
Ideal Profile
Strong candidates may have:
- Advanced practical experience using MCP, plugins, and AI connectors
- Frequent use of connected LLM tools, ideally several times per week
- Heavy personal use of AI for planning, research, organisation, and decision-making
- An active LLM account with approximately 6 months or more of regular usage history
- Experience using AI for high-context, multi-step personal workflows
- Strong written judgment, reasoning, and attention to detail
- Ability to explain clearly why an AI output is effective, incomplete, unsafe, or unrealistic
- Experience designing and applying structured evaluation rubrics
- Availability to contribute at least 20 hours per week
- Ability to complete assigned tasks within approximately 24 hours
- Current residence in the United States
Educational Background
- A degree in computer science, information systems, research, operations, behavioural science, communications, or a related discipline may be helpful
- Formal technical credentials are secondary to extensive hands-on experience with LLM tools and connected workflows
- Professional or project-based experience in AI evaluation, quality assurance, user research, or structured data review may strengthen an application
- Equivalent practical expertise gained through sustained personal and professional AI usage may also be considered
Nice to Have
- More than 100 hours of prior rubric design, evaluation, or quality-assessment experience
- Familiarity with Google Drive, Notion, travel platforms, and other connected applications
- Experience evaluating AI assistants across personal planning and life-organisation tasks
- Background in model evaluation, human data, user research, quality assurance, or AI training
- Strong understanding of context management, personalisation, and tool-use failure modes
- Experience documenting complex workflows through screen recording
- Familiarity with health research, travel planning, dining decisions, home services, or career-search workflows
- Experience identifying subtle issues involving model overreach, missing context, and unrealistic execution
Why This Opportunity
- Help improve how AI assistants support complex real-world personal workflows
- Evaluate advanced systems using practical connected tools and personal context
- Influence how models handle planning, preferences, constraints, and multi-step actions
- Apply deep LLM experience across travel, health, productivity, careers, and everyday decision-making
- Build experience in MCP, connector evaluation, rubric development, and personalised AI research
- Access potential ongoing work following successful completion of the initial trial period
Contract Details
- Independent contractor role
- Fully remote within the United States
- Expected commitment of at least 20 hours per week
- Initial ramp-up period of approximately 1–2 days
- Ability to complete assigned tasks within approximately 24 hours is required
- Desktop or laptop computer required; Chromebooks are not supported
- Screen recording is required during task execution
- Candidates must be willing to sign a data-sharing consent form electronically
- Initial trial period used to assess quality, consistency, and project fit
- Competitive rates between $45–$185 per hour depending on expertise, task complexity, and project scope
- Weekly payments via Stripe or Wise
- Task availability may begin after an initial project setup period
- Projects may be extended, shortened, or adjusted depending on scope and performance
- Work will not involve access to confidential or proprietary information from any employer, client, or institution
About the Platform
This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams.
By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy.