Resilient IT is a security-first managed service provider and cybersecurity compliance firm serving organizations within the Defense Industrial Base. We are an Authorized CMMC Third-Party Assessment Organization (C3PAO), a CMMC Level 2 Certified MSP, and a GTIA Trustmark holder.
We develop and operate our own Governance, Risk, and Compliance platform, along with an internally hosted artificial intelligence platform that is already integrated with portions of the GRC. These technologies support our employees, improve service delivery, strengthen operational consistency, and automate appropriate business and compliance activities.
We are a small but growing organization where employees can take meaningful ownership, solve real problems, and directly influence how the company operates. Our culture is built around security, accountability, continuous improvement, and doing the right thing.
Position Overview
Resilient IT is seeking an experienced, hands-on AI Platform Engineer to assume technical leadership of our existing internally hosted AI platform and advance its integration with our GRC and other internal tools.
This is not a research-only or prompt-engineering position. The AI Engineer will be responsible for making the existing platform reliable, secure, supportable, and increasingly useful in day-to-day operations. This includes managing its technical health, troubleshooting problems, improving model performance, maintaining existing integrations, and developing new AI-assisted capabilities and automations.
Our current platform is built around locally hosted technologies, with supporting components for model routing, retrieval-augmented generation, document processing, vector search, authentication, auditing, and API-based integrations.
The AI Engineer will begin with an established platform and existing GRC integration. The initial responsibility will be to understand the current architecture, validate how its components and integrations operate, resolve outstanding issues, and create a controlled roadmap for improvements.
From there, the role will enhance the platform’s capabilities, expand responsible automation, improve existing workflows, and develop new integrations across ResilientGRC and other approved internal systems.
This Is Not a Typical AI Job
This position requires someone who can move between infrastructure, application development, model behavior, troubleshooting, security, and business process automation.
The right person will not simply deploy models or produce demonstrations. You will take ownership of an operational AI environment, understand what has already been built, identify why components behave the way they do, resolve deficiencies, and improve the platform without disrupting existing functionality.
You will also be expected to recognize when AI should assist a process, when deterministic software is more appropriate, and when a human must remain responsible for the final decision.
What You’ll Do
Assume Ownership of the Existing AI Platform
- Learn and document the current AI architecture, configuration, integrations, workflows, dependencies, and operating procedures.
- Serve as the primary technical owner of Resilient IT’s internally hosted AI platform.
- Manage and optimize Ollama, Qwen models, and the supporting AI application stack.
- Evaluate the current environment for reliability, security, performance, maintainability, and scalability.
- Identify and prioritize existing technical issues, incomplete capabilities, configuration problems, and operational risks.
- Preserve working functionality while implementing controlled improvements.
- Monitor platform availability, model performance, resource utilization, response quality, latency, and system capacity.
- Diagnose and resolve model, infrastructure, application, integration, and data-pipeline issues.
- Manage model installation, configuration, testing, evaluation, upgrades, and controlled deployment.
- Evaluate model configurations, quantization options, context limits, inference settings, and hardware utilization.
- Maintain stable development, testing, and production processes for AI services.
- Establish or improve monitoring, logging, alerting, backup, recovery, and platform health checks.
- Plan for platform scalability, redundancy, and future architectural or hardware improvements.
Improve Model Accuracy and Usefulness
- Review and enhance existing system prompts, structured outputs, reusable workflows, and model-specific configurations.
- Establish evaluation methods for measuring accuracy, completeness, relevance, consistency, hallucination risk, and business value.
- Test model and workflow changes against defined use cases before introducing them into production.
- Develop controlled feedback processes that allow users to report inaccurate or unhelpful results.
- Improve retrieval-augmented generation pipelines, document chunking, metadata, embeddings, vector search, citations, and source grounding.
- Ensure AI-generated answers clearly distinguish authoritative source content from analysis, recommendations, and unresolved factual gaps.
- Reduce unsupported conclusions and prevent the platform from presenting assumptions as verified facts.
- Determine when a larger model, smaller model, specialized workflow, traditional software function, or human review is the appropriate solution.
Enhance the Existing GRC Integration
- Review, maintain, troubleshoot, and enhance the current integration between the AI platform and the GRC.
- Understand the existing APIs, data flows, permissions, AI features, and automation already operating within the platform.
- Preserve current working capabilities while introducing tested and controlled enhancements.
- Resolve reliability, accuracy, performance, usability, and data-handling issues within existing AI-assisted GRC functions.
- Expand AI assistance into additional GRC workflows where there is a defined operational benefit.
- Improve contextual search, guided assistance, document analysis, information extraction, and quality-control capabilities.
- Develop additional secure API-based integrations between the AI platform, GRC, and other approved internal systems.
- Automate repeatable administrative and analytical activities where the resulting output can be validated.
- Ensure all enhancements preserve tenant separation, engagement scope, user permissions, data ownership, auditability, and existing validated functions.
- Make automated actions traceable, reviewable, and reversible when appropriate.
- Maintain clear boundaries between consulting assistance and formal C3PAO assessment activities.
- Ensure AI assists qualified personnel without replacing required assessor judgment, independence, impartiality, or final determinations.
Expand Internal Automation
- Assess existing AI automations and determine where they should be corrected, strengthened, expanded, or retired.
- Work with leadership and department owners to identify additional high-value automation opportunities.
- Translate manual processes into well-defined, controlled, and measurable workflows.
- Integrate the AI platform with approved internal applications, data sources, document repositories, and third-party tools.
- Build agents, services, APIs, background jobs, and workflow automations that perform defined business functions.
- Automate document intake, classification, summarization, comparison, extraction, validation, routing, and reporting where appropriate.
- Develop human-in-the-loop approval processes for sensitive or consequential actions.
- Implement validation rules and exception handling so failed or uncertain automation does not silently produce incorrect results.
- Measure whether each automation saves time, improves consistency, reduces errors, or produces another defined business benefit.
- Maintain an enhancement and automation roadmap based on business priority, technical feasibility, risk, and expected return.
Protect Sensitive Information
- Maintain the platform as a secure, internally controlled environment.
- Apply least privilege, role-based access, authentication, authorization, encryption, logging, and auditability requirements.
- Protect Federal Contract Information, Controlled Unclassified Information, client information, assessment data, and other sensitive records.
- Prevent unauthorized access to information across organizations, tenants, engagements, users, and workflows.
- Ensure prompts, responses, attachments, embeddings, logs, and cached data receive appropriate protection.
- Evaluate new integrations and data flows for privacy, security, compliance, and information-handling risks.
- Identify and remediate prompt injection, data leakage, insecure output handling, excessive agency, and other AI-specific risks.
- Maintain alignment with Resilient IT’s security requirements, internal policies, and applicable customer or regulatory obligations.
Support Users and the Business
- Provide technical support for employees using the AI platform and its GRC capabilities.
- Investigate reported problems and clearly communicate causes, corrective actions, and expected resolution times.
- Develop and maintain user guidance, training materials, operating procedures, and responsible-use standards.
- Help employees understand both the capabilities and limitations of AI-assisted tools.
- Gather user feedback and convert it into prioritized platform improvements.
- Communicate technical decisions in language that business, compliance, and executive stakeholders can understand.
- Participate in planning discussions involving AI capabilities, operational improvements, and product development.
Required Qualifications
- Demonstrated experience operating, administering, enhancing, or developing production AI systems.
- Hands-on experience with locally hosted large language models.
- Experience with Ollama or comparable local model-serving technologies.
- Strong Python development skills.
- Experience designing, consuming, and securing REST APIs.
- Experience with Linux server administration and troubleshooting.
- Experience with Docker and containerized applications.
- Experience assuming ownership of an existing technical platform and improving it without disrupting working functionality.
- Experience integrating AI capabilities into existing business applications.
- Working knowledge of retrieval-augmented generation, embeddings, vector databases, document ingestion, and semantic search.
- Experience designing prompts, structured model outputs, tool-calling workflows, or AI agents.
- Ability to diagnose problems across models, applications, APIs, infrastructure, databases, and networks.
- Understanding of authentication, authorization, secrets management, logging, encryption, and secure application development.
- Ability to evaluate AI output for accuracy, reliability, security, and operational risk.
- Strong documentation, communication, prioritization, and problem-solving skills.
-
Ability to independently own complex systems while collaborating with business and technical stakeholders.
Strongly Preferred
- Experience working directly with Qwen models.
- Experience with Open WebUI, LiteLLM, Qdrant, PostgreSQL, Apache Tika, SearXNG, or similar technologies.
- Experience with FastAPI, SQLModel, React, Vite, or comparable application frameworks.
- Experience managing GPU-based inference, model quantization, memory utilization, and performance optimization.
- Experience with secure, multi-tenant SaaS or internally hosted business platforms.
- Experience with workflow automation, orchestration platforms, background workers, or event-driven integrations.
- Experience implementing model evaluations, regression tests, guardrails, and observability.
- Experience with cybersecurity, compliance, risk management, or regulated environments.
- Familiarity with CMMC, NIST SP 800-171, NIST SP 800-171A, 32 CFR Part 170, or Defense Industrial Base requirements.
- Experience handling Controlled Unclassified Information or other regulated data.
-
Understanding of AI governance, human oversight, explainability, source attribution, and responsible AI practices.
Who You Are
- You can inherit an existing technical environment, understand it, stabilize it, and improve it.
- You take ownership of systems and remain engaged after new capabilities are deployed.
- You enjoy solving difficult technical problems and determining their actual root causes.
- You understand that a convincing AI response is not necessarily a correct response.
- You value reliable, maintainable solutions over impressive but fragile demonstrations.
- You know how to balance experimentation with production stability and security.
- You can turn loosely defined business needs into controlled technical solutions.
- You communicate clearly when a proposed automation is unsafe, unreliable, or inappropriate.
- You document your work so the platform does not depend on undocumented individual knowledge.
- You are comfortable working in a growing organization where priorities evolve and your work has visible business impact.
- You demonstrate integrity, discretion, accountability, and a commitment to doing the right thing.
What Success Looks Like
Success in this role will be measured by:
- Development of a thorough understanding of the existing platform and ResilientGRC integration.
- Effective ownership, maintenance, and documentation of the current environment.
- Reliability, availability, security, and supportability of the internal AI platform.
- Reduction in recurring platform errors, outages, and unresolved technical issues.
- Measurable improvements in model accuracy, response quality, source grounding, and consistency.
- Successful enhancement of existing ResilientGRC AI capabilities without disrupting validated functionality or weakening tenant isolation.
- Delivery of useful new automations that reduce manual effort, shorten turnaround time, or improve quality.
- Increased adoption of the platform by employees for approved business use cases.
- Clear documentation, monitoring, testing, and change-control practices.
- Effective protection of sensitive company and client information.
- Appropriate human review and traceability for AI-assisted decisions and actions.
- Progress against an agreed platform enhancement and automation roadmap.
Schedule and Work Environment
- Full-time, salaried position.
- Standard schedule is Monday through Friday, 8:00 AM to 5:00 PM Eastern Time.
- Primarily remote, with company-provided equipment.
- Occasional on-site work or travel may be required.
- The employee must be available during established business hours to support internal users and respond to platform issues.
- Performance reviews are conducted at approximately 30, 90, and 180 days, followed by annual reviews.
- Compensation is paid semi-monthly.
- Position is eligible for performance-based bonuses.
Compensation
Competitive compensation will be based on experience, technical depth, demonstrated ability, and alignment with the responsibilities of the position.
What We Offer
- Primarily remote work environment.
- Company-provided equipment.
- Paid federal holidays.
- Up to 15 days of paid time off.
- Five personal days.
- Company contribution of 50% toward medical insurance.
- Company-paid individual dental and vision coverage at 99%.
- Company-paid life insurance.
- 401(k) with employer contributions ranging from 1% to 3%.
- Performance-based bonus eligibility.
- Professional development and certification support.
- The opportunity to lead and advance a strategically important internal technology platform.
- Direct involvement in expanding practical AI capabilities for cybersecurity, compliance, and business operations.
- Meaningful growth opportunities as the company, platform, and engineering function expand.
Why Join Resilient IT?
This is an opportunity to take ownership of an established internal AI platform and help move it into its next stage of maturity.
You will not be starting with a blank slate. Resilient IT has already developed its local AI environment and implemented portions of its integration with ResilientGRC. Your responsibility will be to understand what exists, make it work better, resolve issues, expand its capabilities, and use it to deliver secure and measurable automation throughout the company.
You will have a direct role in determining how Resilient IT applies AI across its GRC platform, cybersecurity services, compliance operations, and internal business processes. Your work will influence the platform’s architecture, operating standards, safeguards, integrations, and automation roadmap.
You will join an organization that values security, quality, accountability, innovation, and measurable results. If you want to own a meaningful AI environment, solve challenging operational problems, and see the direct impact of your work, we want to hear from you.
Apply
Applications must be submitted through this site. Please include a current résumé in PDF format and information describing your experience with locally hosted AI models, existing-platform ownership, application integration, automation, Linux, containers, and production AI operations.
Resilient IT is an equal opportunity employer. Employment decisions are based on qualifications, merit, and business needs. We do not discriminate on the basis of any status protected by applicable law.