Skip to content

Latest commit

 

History

8,245 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 


Aruna

A FAIR, federated data orchestration engine

Built with Rust Apache 2.0 license MIT license CI Coverage

Portal documentation · Swagger UI · Features · Architecture · Getting started · Pitfalls · Contributing

Note

Aruna v3 is now in public testing. You can try it out and share feedback through GitHub issues. Aruna v2 remains available on the v2 branch.

Aruna helps research organizations store, describe, share and reuse data while keeping control of their own infrastructure. Research data is often spread across universities, labs, archives and computing centers, each with its own storage systems and access rules. With Aruna, each organization runs a node that connects directly to other nodes. Researchers can find and work with data across institutions, with descriptions that help them understand what each dataset contains, how it was created and how it can be reused.

Features

  • Web portal: browse files, edit datasets, manage access and follow compute runs in the browser.
  • Works with your existing tools: every node offers an S3-compatible interface, so common tools, scripts and workflow systems work without changes.
  • Rich dataset descriptions: datasets are described with RO-Crate, a widely used standard that covers files, people, instruments, software and workflows. Crates reference other crates to connect datasets with their sources and the analyses that produced them.
  • Quality checks: profiles specify which details a dataset description should include, and Aruna highlights anything missing.
  • Search across nodes: find datasets on all connected nodes, limited to what you are allowed to see.
  • Groups and permissions: control who can access your data, with permissions for individuals, groups and roles.
  • Automatic merging: edits made on different nodes are combined automatically when the nodes reconnect, including edits made offline.
  • Native Git for datasets: every dataset is also a Git repository in the ARC layout. Clone it, work on branches and push changes back; the dataset description stays in sync.
  • Dataset history: see every version of a dataset, compare versions and merge draft changes in the portal.
  • Publish with a DOI: send datasets to Zenodo or other InvenioRDM repositories and keep both in sync. Existing records can be imported, too.
  • Compute jobs: run analysis containers on your own infrastructure from the portal or through the standard GA4GH TES interface.
  • Interactive notebooks: write and run Jupyter notebooks in the portal, with direct access to your data.
  • AI assistants: connect AI assistants through MCP so they can search, read and work with your data, with your permissions.
  • Flexible storage: keep data on local disks or connect other storage systems. Buckets can combine local files, copies and references to files on other nodes.
  • Safe transfers: Aruna checks files when storing and copying them to detect data corruption early.
  • Open standards: single sign-on with OIDC, data references with GA4GH DRS and metadata harvesting with OAI-PMH.
  • Simple to deploy: run a single node as one program, or a cluster of nodes.

Architecture and Goals

Institutions need to control where their data is stored and who can access it. Moving files manually takes time, and the context needed to understand them can get lost along the way. Aruna helps institutions share data while retaining control over its storage and access.

  • Nodes: each organization runs its own node. The node decides where the organization's data lives and who may access it.
  • Realms: nodes that trust each other form a realm, for example an institute, a consortium or a project. Being in the same realm does not give anyone access to data; access is always granted explicitly through groups, roles and permissions.
  • Direct connections: nodes can connect to each other automatically, including from behind firewalls. No central server is needed, and a node keeps working when others are offline. Changes are shared again once the nodes can reach each other.
  • Data and description together: files and their descriptions move together, so a dataset stays understandable wherever it is used. Crates reference other crates, forming a provenance graph that connects source data, analyses and results. These links help researchers trace how a result was produced and understand what they need to reproduce it.

The goal is to make research data FAIR: findable, accessible, interoperable and reusable, while each institution stays responsible for its own data.

Getting Started

Start with the public portal, or run Aruna yourself using one of the options below.

1. Try it online

You can explore Aruna without installing anything:

2. Run a small cluster on your computer

This starts three connected nodes on your machine, so you can see how nodes work together. You need docker with Docker Compose and, for convenience, just.

just preview          # three nodes, a login server and the web portal
just local-cluster    # three nodes without the portal

When everything is ready, the command prints the addresses of each node, test logins and an admin token. Press Ctrl-C to stop the cluster, or run just stop if the terminal was closed.

3. Run a single node

To run one node yourself, for example to test it with your own login server:

cp .env.example .env
cargo run -p aruna

Building from source needs the Rust version named in rust-toolchain.toml. The node then serves the API documentation on http://127.0.0.1:3000/swagger-ui and the S3 interface on http://127.0.0.1:1337.

Avoiding Common Pitfalls

Most problems with a new node come from a few settings. These tips help you avoid them.

  • Use your own keys. The example configuration contains demo keys that everybody can see. A node refuses to start with them. Create your own realm and node keys (REALM_*_KEY and NODE_*_KEY) before running a real node.
  • Set up a node right the first time. A node creates or joins its realm only on its very first start: without ONBOARDING_SECRET it creates a new realm, with it it joins an existing one. Changing the configuration afterwards does not move it to another realm. To start over, use a new, empty data directory.
  • Back up the whole data directory. The data directory (STORAGE_PATH) also holds the node's identity, which protects stored secrets such as access tokens. A node restored without it cannot read these secrets anymore.
  • Migrate before upgrading. Stop the node, run aruna-doctor migrate on its data directory, then start the new version. Running the migration twice does no harm.
  • Keep enough free disk space. Importing a large dataset briefly needs about twice its size.
  • Decide how changes are saved. By default, Aruna favors speed: the last few changes can be lost if the machine loses power. Set ARUNA_FJALL_PERSIST_MODE=sync_all if that is not acceptable for you.
  • Keep notebook networks separate. Notebook sessions run in their own network so they can only reach your data. Make sure this network (ARUNA_COMPUTE_DOCKER_SESSION_SUBNET) does not overlap with other networks on the host.
  • Use Git the usual way. Log in to the Git repository of a dataset with an Aruna access token as the password, pull before you push, and store large files with Git LFS.

Learn More

License

Aruna is licensed under either of

at your option. Unless you explicitly state otherwise, any contribution intentionally submitted for inclusion in Aruna by you, as defined in the Apache-2.0 license, shall be dual licensed as above, without any additional terms or conditions.

Feedback & Contributions

Found a bug or have an idea? Open an issue or send a pull request. Reports from the v3 public test phase help us understand what works and what needs attention. See the Contributor Guidelines and Code of Conduct before contributing.

About

The data orchestration engine

Resources

Code of conduct

Contributing

Stars

24 stars

Watchers

2 watching

Forks

Releases

Packages

Used by

Contributors

Languages