Research Scientist – Computer Vision

Scientist

Research Scientist – Computer Vision

Apply Now

- $0.00

  • Date posted
    August 10, 2026
  • Expiration date
    November 10, 2026
  • Application ends
    November 10, 2026

We are seeking a Research Scientist specialising in computer vision and multimodal AI to join a leading research and development team

 

Key Responsibilities

Advanced Model Research

  • Design and develop Vision Transformer and multimodal large-model architectures with improved reasoning, efficiency, and scalability.
  • Advance multimodal alignment, representation learning, and long-context modelling.
  • Research scalable training techniques for large multimodal models.
  • Improve model architectures to strengthen generalisation, robustness, and overall performance.
  • Explore new approaches across multimodal understanding, generation, and reasoning.

Multimodal Data Development

  • Process large-scale multimodal datasets spanning images, video, audio, and text.
  • Build data pipelines for cleaning, filtering, annotation, validation, and quality control.
  • Construct and maintain reproducible datasets with clear versioning and documentation.
  • Optimise data mixtures, sampling methods, and curriculum strategies for model training.
  • Improve dataset quality through evaluation results and feedback-driven curation.
  • Develop methods for identifying low-quality, duplicated, biased, or uninformative data.

Large Multimodal Model Systems

  • Build and improve distributed training systems for large-scale multimodal models.
  • Optimise GPU utilisation, cluster efficiency, resource allocation, and workload scheduling.
  • Develop scalable training frameworks and reusable research infrastructure.
  • Engineer training, inference, evaluation, and serving systems.
  • Improve the scalability, reliability, stability, and performance of model development pipelines.
  • Work closely with infrastructure and platform teams to resolve system-level bottlenecks.

Research Application and Delivery

  • Apply multimodal capabilities to intelligent assistants, content generation, and related AI applications.
  • Translate research outcomes into production-ready systems and user-facing features.
  • Collaborate with product, engineering, and research teams to deploy, evaluate, and iterate models.
  • Communicate research findings through technical reports, presentations, publications, and demonstrations.

Essential Requirements

  • Bachelor’s degree or higher in Computer Science, Mathematics, Statistics, Engineering, or another relevant technical discipline.
  • Strong Python programming skills.
  • Hands-on experience with PyTorch or comparable deep learning frameworks.
  • Strong ability to develop, implement, and evaluate machine learning algorithms.
  • Solid mathematical reasoning and problem-solving skills.
  • Good understanding of modern computer vision or multimodal machine learning methods.
  • Ability to collaborate effectively across research, engineering, product, and infrastructure teams.
  • Strong written and verbal communication skills.
  • Self-motivated, resilient, and comfortable working on complex research problems with a high degree of technical uncertainty.

Preferred Qualifications

  • Master’s degree or PhD in Computer Science, Artificial Intelligence, Machine Learning, Computer Vision, or a related field.
  • Publications at leading artificial intelligence or computer vision conferences, including CVPR, ICCV, ECCV, NeurIPS, ICML, or ICLR.
  • Experience pre-training, fine-tuning, or evaluating large-scale vision or multimodal models.
  • Experience with Vision Transformers, vision-language models, multimodal large language models, or generative vision systems.
  • Familiarity with distributed training, mixed-precision training, model parallelism, or large-scale GPU clusters.
  • Contributions to high-impact open-source projects in computer vision, natural language processing, multimodal AI, or machine learning systems.
  • Research or internship experience within a recognised technology company, research laboratory, or academic institution.
  • Experience translating research prototypes into scalable production systems.

Additional Skills

  • Strong understanding of representation learning, attention mechanisms, transformers, and generative modelling.
  • Experience working with large multimodal datasets and data-quality pipelines.
  • Familiarity with model evaluation, benchmarking, ablation studies, and experiment reproducibility.
  • Ability to identify research opportunities and independently drive projects from initial concept to validated outcome.
  • Interest in advancing the capabilities, efficiency, and reliability of next-generation multimodal AI systems.
  • Are you interested in this position?

     

    Apply by clicking on the “Apply Now” button below!

     

    #AlbionarcJobs#FintechJobs

    #AsiaJobs#MiddleEastCareers

    #TechTalent#FintechRecruitment

    #FinanceOpportunities#

     

Apply Now

- $0.00

Select your currency