![]() |
WRENCH
1.11
Cyberinfrastructure Simulation Workbench
|
Overview | Installation | Getting Started | WRENCH 101 | WRENCH 102 |
|
A compute service that manages a set of multi-core compute hosts and provides access to their resources. More...
#include <BareMetalComputeService.h>
Public Member Functions | |
BareMetalComputeService (const std::string &hostname, const std::map< std::string, std::tuple< unsigned long, double >> compute_resources, std::string scratch_space_mount_point, WRENCH_PROPERTY_COLLECTION_TYPE property_list={}, WRENCH_MESSAGE_PAYLOADCOLLECTION_TYPE messagepayload_list={}) | |
Constructor. More... | |
BareMetalComputeService (const std::string &hostname, const std::vector< std::string > compute_hosts, std::string scratch_space_mount_point, WRENCH_PROPERTY_COLLECTION_TYPE property_list={}, WRENCH_MESSAGE_PAYLOADCOLLECTION_TYPE messagepayload_list={}) | |
Constructor. More... | |
~BareMetalComputeService () | |
Destructor. | |
void | submitCompoundJob (std::shared_ptr< CompoundJob > job, const std::map< std::string, std::string > &service_specific_args) override |
Submit a compound job to the compute service. More... | |
virtual bool | supportsCompoundJobs () override |
Returns true if the service supports compound jobs. More... | |
virtual bool | supportsPilotJobs () override |
Returns true if the service supports pilot jobs. More... | |
virtual bool | supportsStandardJobs () override |
Returns true if the service supports standard jobs. More... | |
void | terminateCompoundJob (std::shared_ptr< CompoundJob > job) override |
Synchronously terminate a compound job previously submitted to the compute service. More... | |
![]() | |
ComputeService (const std::string &hostname, std::string service_name, std::string scratch_space_mount_point) | |
Constructor. More... | |
std::map< std::string, double > | getCoreFlopRate () |
Get the per-core flop rate of the compute service's hosts. More... | |
double | getFreeScratchSpaceSize () |
Get the free space on the compute service's scratch storage space. More... | |
std::vector< std::string > | getHosts () |
Get the list of the compute service's compute host. More... | |
std::map< std::string, double > | getMemoryCapacity () |
Get the RAM capacities for each of the compute service's hosts. More... | |
unsigned long | getNumHosts () |
Get the number of hosts that the compute service manages. More... | |
std::map< std::string, double > | getPerHostAvailableMemoryCapacity () |
Get ram availability for each of the compute service's host. More... | |
std::map< std::string, unsigned long > | getPerHostNumCores () |
Get core counts for each of the compute service's host. More... | |
std::map< std::string, unsigned long > | getPerHostNumIdleCores () |
Get idle core counts for each of the compute service's host. More... | |
std::shared_ptr< StorageService > | getScratch () |
Get the compute service's scratch storage space. More... | |
unsigned long | getTotalNumCores () |
Get the total core counts for all hosts of the compute service. More... | |
virtual unsigned long | getTotalNumIdleCores () |
Get the total idle core count for all hosts of the compute service. Note that this doesn't mean that asking for these cores right will mean immediate execution (since jobs may be pending and "ahead" in the queue, e.g., because they depend on current actions that are not using all available resources). More... | |
double | getTotalScratchSpaceSize () |
Get the total capacity of the compute service's scratch storage space. More... | |
double | getTTL () |
Get the time-to-live of the compute service. More... | |
virtual bool | hasScratch () const |
Checks if the compute service has a scratch space. More... | |
virtual bool | isThereAtLeastOneHostWithIdleResources (unsigned long num_cores, double ram) |
Method to find out if, right now, the compute service has at least one host with some idle number of cores and some available RAM. Note that this doesn't mean that asking for these resources right will mean immediate execution (since jobs may be pending and "ahead" in the queue, e.g., because they depend on current actions that are not using all available resources). More... | |
void | stop () override |
Stop the compute service. | |
virtual void | stop (bool send_failure_notifications, ComputeService::TerminationCause termination_cause) |
Stop the compute service. More... | |
void | terminateJob (std::shared_ptr< CompoundJob > job) |
Terminate a previously-submitted job (which may or may not be running yet) More... | |
![]() | |
void | assertServiceIsUp () |
Throws an exception if the service is not up. More... | |
std::string | getHostname () |
Get the name of the host on which the service is / will be running. More... | |
const WRENCH_MESSAGE_PAYLOADCOLLECTION_TYPE & | getMessagePayloadList () const |
Get all message payloads and their values of the Service. More... | |
double | getMessagePayloadValue (WRENCH_MESSAGEPAYLOAD_TYPE) |
Get a message payload of the Service as a double. More... | |
double | getNetworkTimeoutValue () |
Returns the service's network timeout value. More... | |
std::string | getPhysicalHostname () |
Get the physical name of the host on which the service is / will be running. More... | |
bool | getPropertyValueAsBoolean (WRENCH_PROPERTY_TYPE) |
Get a property of the Service as a boolean. More... | |
double | getPropertyValueAsDouble (WRENCH_PROPERTY_TYPE) |
Get a property of the Service as a double. More... | |
std::string | getPropertyValueAsString (WRENCH_PROPERTY_TYPE) |
Get a property of the Service as a string. More... | |
unsigned long | getPropertyValueAsUnsignedLong (WRENCH_PROPERTY_TYPE) |
Get a property of the Service as an unsigned long. More... | |
bool | isUp () |
Returns true if the service is UP, false otherwise. More... | |
void | resume () |
Resume the service. More... | |
void | setNetworkTimeoutValue (double value) |
Sets the service's network timeout value. More... | |
void | setStateToDown () |
Set the state of the service to DOWN. | |
void | start (std::shared_ptr< Service > this_service, bool daemonize, bool auto_restart) |
Start the service. More... | |
void | suspend () |
Suspend the service. | |
![]() | |
S4U_Daemon (std::string hostname, std::string process_name_prefix) | |
Constructor (daemon with a mailbox) More... | |
void | acquireDaemonLock () |
Method to acquire the daemon's lock. More... | |
void | createLifeSaver (std::shared_ptr< S4U_Daemon > reference) |
Create a life saver for the daemon. More... | |
std::string | getName () |
Retrieve the process name. More... | |
int | getReturnValue () |
Returns the value returned by main() (if the daemon has returned from main) More... | |
Simulation * | getSimulation () |
Get the service's simulation. More... | |
S4U_Daemon::State | getState () |
Get the daemon's state. More... | |
bool | hasReturnedFromMain () |
Returns true if the daemon has returned from main() (i.e., not brutally killed) More... | |
bool | isDaemonized () |
Return the daemonized status of the daemon. More... | |
bool | isSetToAutoRestart () |
Return the auto-restart status of the daemon. More... | |
std::pair< bool, int > | join () |
Join (i.e., wait for) the daemon. More... | |
void | releaseDaemonLock () |
Method to release the daemon's lock. More... | |
void | resumeActor () |
Resume the daemon/actor. | |
void | setSimulation (Simulation *simulation) |
Set the service's simulation. More... | |
void | setupOnExitFunction () |
Sets up the on_exit function for the actor. | |
void | startDaemon (bool _daemonized, bool _auto_restart) |
Start the daemon. More... | |
void | suspendActor () |
Suspend the daemon/actor. | |
Protected Member Functions | |
BareMetalComputeService (const std::string &hostname, std::map< std::string, std::tuple< unsigned long, double >> compute_resources, WRENCH_PROPERTY_COLLECTION_TYPE property_list, WRENCH_MESSAGE_PAYLOADCOLLECTION_TYPE messagepayload_list, double ttl, std::shared_ptr< PilotJob > pj, std::string suffix, std::shared_ptr< StorageService > scratch_space) | |
Internal constructor. More... | |
BareMetalComputeService (const std::string &hostname, std::map< std::string, std::tuple< unsigned long, double >> compute_resources, WRENCH_PROPERTY_COLLECTION_TYPE property_list, WRENCH_MESSAGE_PAYLOADCOLLECTION_TYPE messagepayload_list, std::shared_ptr< StorageService > scratch_space) | |
Internal constructor. More... | |
void | cleanup (bool has_terminated_cleanly, int return_value) override |
Cleanup method. More... | |
void | cleanUpScratch () |
Cleans up the scratch as I am a pilot job and I to need clean the files stored by the standard jobs executed inside me. | |
void | dispatchReadyActions () |
Helper method to dispatch actions. | |
void | initiateInstance (const std::string &hostname, std::map< std::string, std::tuple< unsigned long, double >> compute_resources, WRENCH_PROPERTY_COLLECTION_TYPE property_list, WRENCH_MESSAGE_PAYLOADCOLLECTION_TYPE messagepayload_list, double ttl, std::shared_ptr< PilotJob > pj) |
Helper method called by all constructors to initiate object instance. More... | |
int | main () override |
Main method of the daemon. More... | |
void | processActionDone (std::shared_ptr< Action > action) |
Process an action completion. More... | |
void | processCompoundJobTerminationRequest (std::shared_ptr< CompoundJob > job, simgrid::s4u::Mailbox *answer_mailbox) |
Process a compound job termination request. More... | |
void | processGetResourceInformation (simgrid::s4u::Mailbox *answer_mailbox, const std::string &key) |
Process a "get resource description message". More... | |
void | processIsThereAtLeastOneHostWithAvailableResources (simgrid::s4u::Mailbox *answer_mailbox, unsigned long num_cores, double ram) |
Process a host available resource request. More... | |
bool | processNextMessage () |
Wait for and react to any incoming message x. More... | |
void | processSubmitCompoundJob (simgrid::s4u::Mailbox *answer_mailbox, std::shared_ptr< CompoundJob > job, std::map< std::string, std::string > &service_specific_arguments) |
Process a submit compound job request. More... | |
void | terminate (bool send_failure_notifications, ComputeService::TerminationCause termination_cause) |
Terminate the daemon, dealing with pending/running job. More... | |
void | terminateCurrentCompoundJob (std::shared_ptr< CompoundJob > job, ComputeService::TerminationCause termination_cause) |
Method for terminating a current compound job. More... | |
void | validateProperties () |
Method to make sure that property specs are valid. More... | |
void | validateServiceSpecificArguments (std::shared_ptr< CompoundJob > job, std::map< std::string, std::string > &service_specific_args) override |
Method the validates service-specific arguments (throws std::invalid_argument if invalid) More... | |
![]() | |
ComputeService (const std::string &hostname, std::string service_name, std::shared_ptr< StorageService > scratch_space) | |
Constructor. More... | |
std::shared_ptr< StorageService > | getScratchSharedPtr () |
Get a shared pointer to the compute service's scratch storage space. More... | |
void | submitJob (std::shared_ptr< CompoundJob > job, const std::map< std::string, std::string > &={}) |
Submit a job to the compute service. More... | |
virtual void | validateJobsUseOfScratch (std::map< std::string, std::string > &service_specific_args) |
Method to validate that a job's use of the scratch space if ok. Throws exception if not. More... | |
![]() | |
Service (std::string hostname, std::string process_name_prefix) | |
Constructor. More... | |
~Service () override | |
Destructor. | |
template<class T > | |
std::shared_ptr< T > | getSharedPtr () |
Method to retrieve the shared_ptr to a service. More... | |
void | serviceSanityCheck () |
Check whether the service is properly configured and running. More... | |
void | setMessagePayload (WRENCH_MESSAGEPAYLOAD_TYPE, double) |
Set a message payload of the Service. More... | |
void | setMessagePayloads (WRENCH_MESSAGE_PAYLOADCOLLECTION_TYPE default_messagepayload_values, WRENCH_MESSAGE_PAYLOADCOLLECTION_TYPE overriden_messagepayload_values) |
Set default and user-defined message payloads. More... | |
void | setProperties (WRENCH_PROPERTY_COLLECTION_TYPE default_property_values, WRENCH_PROPERTY_COLLECTION_TYPE overriden_property_values) |
Set default and user-defined properties. More... | |
void | setProperty (WRENCH_PROPERTY_TYPE, const std::string &) |
Set a property of the Service. More... | |
![]() | |
bool | killActor () |
Kill the daemon/actor (does nothing if already dead) More... | |
void | runMainMethod () |
Method that run's the user-defined main method (that's called by the S4U actor class) | |
Static Protected Member Functions | |
static std::tuple< std::string, unsigned long > | parseResourceSpec (const std::string &spec) |
Helper static method to parse resource specifications to the <cores,ram> format. More... | |
![]() | |
static void | assertServiceIsUp (std::shared_ptr< Service > s) |
Assert for the service being up. More... | |
Protected Attributes | |
std::shared_ptr< ActionExecutionService > | action_execution_service |
std::shared_ptr< PilotJob > | containing_pilot_job |
std::set< std::shared_ptr< CompoundJob > > | current_jobs |
std::shared_ptr< Alarm > | death_alarm = nullptr |
double | death_date |
std::set< std::shared_ptr< Action > > | dispatched_actions |
int | exit_code = 0 |
std::unordered_map< std::shared_ptr< StandardJob >, std::set< std::shared_ptr< DataFile > > > | files_in_scratch |
bool | has_ttl |
std::shared_ptr< HostStateChangeDetector > | host_state_change_monitor |
std::set< std::shared_ptr< Action > > | not_ready_actions |
std::unordered_map< std::shared_ptr< CompoundJob >, int > | num_dispatched_actions_for_cjob |
std::vector< std::shared_ptr< Action > > | ready_actions |
double | ttl |
![]() | |
std::shared_ptr< StorageService > | scratch_space_storage_service |
A scratch storage service associated to the compute service. | |
![]() | |
WRENCH_MESSAGE_PAYLOADCOLLECTION_TYPE | messagepayload_list |
The service's messagepayload list. | |
std::string | name |
The service's name. | |
double | network_timeout = 30.0 |
The time (in seconds) after which a service that doesn't send back a reply (control) message causes a NetworkTimeOut exception. (default: 30 second; if <0 never timeout) | |
WRENCH_PROPERTY_COLLECTION_TYPE | property_list |
The service's property list. | |
bool | shutting_down = false |
A boolean that indicates if the service is in the middle of shutting down. | |
![]() | |
unsigned int | num_starts = 0 |
The number of time that this daemon has started (i.e., 1 + number of restarts) | |
Simulation * | simulation |
a pointer to the simulation object | |
State | state |
The service's state. | |
Additional Inherited Members | |
![]() | |
enum | TerminationCause { TERMINATION_NONE, TERMINATION_COMPUTE_SERVICE_TERMINATED, TERMINATION_JOB_KILLED, TERMINATION_JOB_TIMEOUT } |
Job termination cause enum. | |
![]() | |
enum | State { UP, DOWN, SUSPENDED } |
Daemon states. More... | |
![]() | |
static simgrid::s4u::Mailbox * | getRunningActorRecvMailbox () |
Return the running actor's recv mailbox. More... | |
![]() | |
std::string | hostname |
The name of the host on which the daemon is running. | |
LifeSaver * | life_saver = nullptr |
The daemon's life saver. | |
simgrid::s4u::Mailbox * | mailbox |
The daemon's mailbox. | |
std::string | process_name |
The name of the daemon. | |
simgrid::s4u::Mailbox * | recv_mailbox |
The daemon's receive mailbox (to send to another daemon so that that daemon can reply) | |
![]() | |
static constexpr unsigned long | ALL_CORES = ULONG_MAX |
A convenient constant to mean "use all cores of a physical host" whenever a number of cores is needed when instantiating compute services. | |
static constexpr double | ALL_RAM = DBL_MAX |
A convenient constant to mean "use all ram of a physical host" whenever a ram capacity is needed when instantiating compute services. | |
![]() | |
static std::unordered_map< aid_t, simgrid::s4u::Mailbox * > | map_actor_to_recv_mailbox |
A compute service that manages a set of multi-core compute hosts and provides access to their resources.
One can think of this as a simple service that allows the user to run tasks and to specify for each task on which host it should run and with how many cores. If no host is specified, the service will pick the least loaded host. If no number of cores is specified, the service will use as many cores as possible. The service will make sure that the RAM capacity of a host is not exceeded by possibly delaying task executions until enough RAM is available.
wrench::BareMetalComputeService::BareMetalComputeService | ( | const std::string & | hostname, |
const std::map< std::string, std::tuple< unsigned long, double >> | compute_resources, | ||
std::string | scratch_space_mount_point, | ||
WRENCH_PROPERTY_COLLECTION_TYPE | property_list = {} , |
||
WRENCH_MESSAGE_PAYLOADCOLLECTION_TYPE | messagepayload_list = {} |
||
) |
Constructor.
hostname | the name of the host on which the service should be started |
compute_resources | a map of <num_cores, memory_manager_service> tuples, indexed by hostname, which represents the compute resources available to this service.
|
scratch_space_mount_point | the compute service's scratch space's mount point ("" means none) |
property_list | a property list ({} means "use all defaults") |
messagepayload_list | a message payload list ({} means "use all defaults") |
wrench::BareMetalComputeService::BareMetalComputeService | ( | const std::string & | hostname, |
const std::vector< std::string > | compute_hosts, | ||
std::string | scratch_space_mount_point, | ||
WRENCH_PROPERTY_COLLECTION_TYPE | property_list = {} , |
||
WRENCH_MESSAGE_PAYLOADCOLLECTION_TYPE | messagepayload_list = {} |
||
) |
Constructor.
hostname | the name of the host on which the service should be started |
compute_hosts | the names of the hosts available as compute resources (the service will use all the cores and all the RAM of each host) |
scratch_space_mount_point | the compute service's scratch space's mount point ("" means none) |
property_list | a property list ({} means "use all defaults") |
messagepayload_list | a message payload list ({} means "use all defaults") |
|
protected |
Internal constructor.
hostname | the name of the host on which the service should be started |
compute_resources | a list of <hostname, num_cores, memory_manager_service> tuples, which represent the compute resources available to this service |
property_list | a property list ({} means "use all defaults") |
messagepayload_list | a message payload list ({} means "use all defaults") |
ttl | the time-to-live, in seconds (DBL_MAX: infinite time-to-live) |
pj | a containing PilotJob (nullptr if none) |
suffix | a string to append to the process name |
scratch_space | the scratch storage service |
std::invalid_argument |
|
protected |
Internal constructor.
hostname | the name of the host on which the job executor should be started |
compute_resources,: | a list of <hostname, num_cores, memory_manager_service> tuples, which represent the compute resources available to this service |
property_list | a property list ({} means "use all defaults") |
messagepayload_list | a message payload list ({} means "use all defaults") |
scratch_space | the scratch space for this compute service |
|
overrideprotectedvirtual |
Cleanup method.
has_returned_from_main | whether main() returned |
return_value | the return value (if main() returned) |
Reimplemented from wrench::S4U_Daemon.
|
protected |
Helper method called by all constructors to initiate object instance.
hostname | the name of the host |
compute_resources | compute_resources: a map of <num_cores, memory_manager_service> pairs, indexed by hostname, which represent the compute resources available to this service |
property_list | a property list ({} means "use all defaults") |
messagepayload_list | a message payload list ({} means "use all defaults") |
ttl | the time-to-live, in seconds (DBL_MAX: infinite time-to-live) |
pj | a containing PilotJob (nullptr if none) |
std::invalid_argument |
|
overrideprotectedvirtual |
|
staticprotected |
Helper static method to parse resource specifications to the <cores,ram> format.
spec | specification string |
std::invalid_argument |
|
protected |
Process an action completion.
action |
|
protected |
Process a compound job termination request.
job | the job to terminate |
answer_mailbox | the mailbox to which the answer message should be sent |
|
protected |
Process a "get resource description message".
answer_mailbox | the mailbox to which the description message should be sent |
key | the desired resource information (i.e., dictionary key) that's needed) |
|
protected |
Process a host available resource request.
answer_mailbox | the answer mailbox |
num_cores | the desired number of cores |
ram | the desired RAM |
|
protected |
Wait for and react to any incoming message x.
std::runtime_error |
|
protected |
Process a submit compound job request.
answer_mailbox | the mailbox to which the answer message should be sent |
job | the job |
service_specific_arguments | service specific arguments |
|
overridevirtual |
Submit a compound job to the compute service.
job | a compound job |
service_specific_args | optional service specific arguments |
These arguments are provided as a map of strings, indexed by action names. These strings are formatted as "[hostname:][num_cores]" (e.g., "somehost:12", "somehost","6", "").
ExecutionException | |
std::invalid_argument | |
std::runtime_error |
Implements wrench::ComputeService.
|
overridevirtual |
Returns true if the service supports compound jobs.
Implements wrench::ComputeService.
|
overridevirtual |
Returns true if the service supports pilot jobs.
Implements wrench::ComputeService.
|
overridevirtual |
Returns true if the service supports standard jobs.
Implements wrench::ComputeService.
|
protected |
Terminate the daemon, dealing with pending/running job.
send_failure_notifications | whether to send failure notifications |
termination_cause | termination cause (if failure notifications are sent) |
|
overridevirtual |
Synchronously terminate a compound job previously submitted to the compute service.
job | a compound job |
ExecutionException | |
std::runtime_error |
Implements wrench::ComputeService.
|
protected |
Method for terminating a current compound job.
job | the job |
termination_cause | the sermination cause |
|
protected |
Method to make sure that property specs are valid.
std::invalid_argument |
|
overrideprotectedvirtual |
Method the validates service-specific arguments (throws std::invalid_argument if invalid)
job | the job that's being submitted |
service_specific_args | the service-specific arguments |
Reimplemented from wrench::ComputeService.