Computational Scientists

# Cloud-based, collaborative bioinformatics for computational research

Code Ocean makes it easy for computational scientists and bioinformaticians to perform analysis, build pipelines, and reproduce their results in a secure cloud-based environment.

## Key benefits for computational scientists

### Get started in the cloud in minutes

Code Ocean makes it easy to work in the cloud. Attach data from your S3 buckets to any Compute Capsule or Pipeline, assign a flex or dedicated compute resource, then start development in your preferred IDE and any open source language like R or Python.

### Build pipelines in Nextflow without writing code

Build in a visual editor while Code Ocean writes Nextflow in the background in real-time. Import pre-existing pipelines directly from nf-core, clone from a Git repo, or unlock to write your own Nextflow. Runs on a preconfigured AWS Batch instance to parallelize compute and scale storage with EBS autoscaling.

### Collaborate with other colleagues and teams

Share everything from basic analyses to complex pipelines with individual users or whole groups within your organization. Assign permissions to co-develop, or let them duplicate and take things in a new direction.

### Import existing work from your Git provider

Code Ocean connects to multiple different git providers. New Capsules and Pipelines in Code Ocean can be created by duplicating code from private or public git sources to run within your Code Ocean deployment.

### Automate result provenance & traceability

All Result Data include a lineage graph showing data sources and processing steps. All runs are captured in the Timeline and all development is tracked by Git, which can be optionally synced with your preferred provider.

## What organizations use Code Ocean?

Biotechs

Code Ocean is for biotechs starting out, expanding, or working at scale. Get guaranteed reproducible computational science in the cloud from day one, without the DevOps workload.

Pharma companies

Code Ocean is for pharmaceutical companies that value reliable compliance, collaboration, and reproducibility in cloud-based multidisciplinary computational research teams.

Research institutes

Code Ocean is for research institutes that want to improve computational collaboration, connect and work with multiple different sources of data, and improve reproducibility.

## Key product features

Compute Capsules

Code, data, and environment packaged in a single object.

Pipelines

Connect, automate, parallelize, and scale computational work.

Models

Integrated MLflow for tracking, registration, lineage, & inference.

Data

Manage all your Data assets in the cloud and  from multiple sources.

Lineage Graph

An immutable record of how Result Data are generated.

Collections

Gather and organize your work to make it visible and accessible.

Apps

Use no-code bioinformatics apps, or build new ones from scratch.

Admin

A single pane of glass for the entire computational R&D cloud.

## Built for Computational Science

### Data analysis

Use ready-made template Compute Capsules to analyze your data, develop your data analysis workflow in your preferred language and IDE using any open-source software, and take advantage of built-in containerization to guarantee reproducibility.

### Data management

Manage your organization's data and control who has access to it. Built specifically to meet all FAIR principles, data management in Code Ocean uses custom metadata and controlled vocabularies to ensure consistency and improve searchability.

### Bioinformatics pipelines

Build, configure and monitor bioinformatics pipelines from scratch using a visual builder for easy set-up. Or, import from nf-core in one click for instant access to a curated set of best practice analysis pipelines. Runs on AWS Batch out-of-the-box, so your pipelines scale automatically. No setup needed.

### ML model development

Code Ocean is uniquely suited for Artificial Intelligence, Machine Learning, Deep Learning, and Generative AI. Install GPU-ready environments and provision GPU resources in a few clicks. Integration with MLFlow allows you to develop models, track parameters, manage models from development to production, while enjoying out-of-the-box reproducibility and lineage.

### Multiomics

Analyze and work with large multimodal datasets efficiently using scalable compute and storage resources, cached packages for R and Python, preloaded multiomics analysis software that works out of the box and full lineage and reproducibility.

### Imaging

Process images using a variety of tools: from dedicated desktop applications to custom-written deep learning pipelines, from a few individual files to petabyte-sized datasets. No DevOps required, always with lineage.

### Cloud management

Code Ocean makes it easy to manage data and provision compute: CPUs, GPUs, and RAM. Assign flex machines and dedicated machines to manage what is available to your users. Spot instances, idleness detection, and automated shutdown help reduce cloud costs.

### Data/model provenance

Keep track of all data and results with automated result provenance and lineage graph generation. Assess reproducibility with a visual representation of every Capsule, Pipeline, and Data asset involved in a computation.
